scieee AI-readable full text Open interactive document viewer

Structured 3D features for reconstructing controllable avatars

Corona Puyane, Enric,Zanfir, Mihai,Alldieck, Thiemo,Gabriel Bazavan, Eduard,Zanfir, Andrei,Sminchisescu, Cristian

Abstract

We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for optimal coverage of the person of interest, beyond just the body shape, which in turn, additionally helps modeling accessories, hair, and loose clothing. Owing to this, we present a complete 3D transformer-based attention framework which, given a single image of a person in an unconstrained pose, generates an animatable 3D reconstruction with albedo and illumination decomposition, as a result of a single end-to-end model, trained semi-supervised, and with no additional postprocessing. We show that our S3F model surpasses the previous state-of-the-art on various tasks, including monocular 3D reconstruction, as well as albedo & shading estimation. Moreover, we show that the proposed methodology allows novel view synthesis, relighting, and re-posing the reconstruction, and can naturally be extended to handle multiple input images (e.g. different views of a person, or the same view, in different poses, in video). Finally, we demonstrate the editing capabilities of our model for 3D virtual try-on applications.

Full text

Structured 3D Features for Reconstructing Controllable Avatars Enric Corona1†Mihai Zanfir2†Thiemo Alldieck3 Eduard Gabriel Bazavan3Andrei Zanfir3Cristian Sminchisescu3 1UPC, Barcelona 2Newton 3Google Research New viewNew Illumination New poseAlbedo ShadingGeometry 3D Reconstruction Monocular image New clothing View synthesis and 3D editing Figure 1. We introduce Structured 3D Features (S3F), a new feature representation for monocular 3D Human Reconstruction. Using S3Fs, we train a new implicit reconstruction model for creating rigged and relightable avatars – without the need for post-processing. S3Fs store features in a dense 3D point cloud around a human body model, allowing for various applications: our model can produce 3D reconstructions in novel poses, under new illumination, under novel views, or with altered clothing for virtual try-on. Abstract We introduce Structured 3D Features, a model based on a novel implicit 3D representation that pools pixel-aligned image features onto dense 3D points sampled from a parametric, statistical human mesh surface. The 3D points have associated semantics and can move freely in 3D space. This allows for optimal coverage of the person of interest, beyond just the body shape, which in turn, additionally helps modeling accessories, hair, and loose clothing. Owing to this, we present a complete 3D transformer-based attention framework which, given a single image of a person in an unconstrained pose, generates an animatable 3D reconstruction with albedo and illumination decomposition, as a result of a single end-to-end model, trained semi-supervised, and with no additional postprocessing. We show that our S3F model surpasses the previous state-of-the-art on various tasks, including monocular 3D reconstruction, as well as albedo & shading estimation. Moreover, we show that the proposed methodology allows novel view synthesis, relighting, and re-posing the reconstruction, and can naturally be extended to handle multiple input images (e.g. different views of a person, or the same view, in different poses, in video). Finally, we demonstrate the editing capabilities of our model for 3D virtual try-on applications. †Work was done while Enric and Mihai were with Google Research. 1. Introduction Human digitization is playing a major role in several important applications, including AR/VR, video games, social telepresence, virtual try-on, or the movie industry. Traditionally, 3D virtual avatars have been created using multiview stereo [39] or expensive equipment [56]. More flexible or low-cost solutions are often template-based [1,4], but these lack expressiveness for representing details such as hair. Recently, research on implicit representations [6, 17,26,61,62] and neural fields [29,53,72,81] has made significant progress in improving the realism of avatars.The proposed models produce detailed results and are also capable, within limits, to represent loose hair or clothing. Even though model-free methods yield high-fidelity avatars, they are not suited for downstream tasks such as animation. Furthermore, proficiency is often limited to a certain range of imaged body poses. Aiming at these problems, different efforts have been made to combine parametric models with more flexible implicit representations [25,26,71,82]. These methods support animation [25,26] or tackle challenging body poses [71,82]. However, most work relies on 2D pixel-aligned features. This leads to two important problems: (1) First, errors in body pose or camera parameter estimation will result in misalignment between the projections of 3D points on the body surface and image features, which ultimately results in low quality reconstructions. (2) image features are coupled with arXiv:2212.06820v2 [cs.CV] 24 Mar 2023 the input view and cannot be easily manipulated e.g. to edit the reconstructed body pose. In this paper we introduce Structure 3D Features (S3F), a flexible extension to image features, specifically designed to tackle the previously discussed challenges and to provide more flexibility during and after digitization. S3F store local features on ordered sets of points around the body surface, taking advantage of the geometric body prior. As body models do not usually represent hair or loose clothing, it is difficult to recover accurate body parameters for images inthe-wild. To this end, instead of relying too much on the geometric body prior, our model freely moves 3D body points independently to cover areas that are not well represented by the prior. This process results in our novel S3Fs, and is trained without explicit supervision only using reconstruction losses as signals. Another limitation of prior work is its dependence on 3D scans. We alleviate this dependence by following a mixed strategy: we combine a small collection of 3D scans, typically the only training data considered in previous work, with large-scale monocular inthe-wild image collections. We show that by guiding the training process with a small set of 3D synthetic scans, the method can efficiently learn features that are only available for the scans (e.g. albedo), while self-supervision on real images allow the method to generalize better to diverse appearances, clothing types and challenging body poses. In this paper, we show how the proposed S3Fs are substantially more flexible than current state-of-the-art representations. S3Fs enable us to train a single end-to-end model that, based on an input image and matching body pose parameters, can generate a 3D human reconstruction that is relightable and animatable. Furthermore, our model supports the 3D editing of e.g. clothing, without additional post-processing. See Fig.1for an illustration and Table 1 for a summary of our model’s properties, in relation to prior work. We compare our method with prior work and demonstrate state-of-the-art performance for monocular 3D human reconstruction from challenging, real-world and unconstrained images, and for albedo & shading estimation. We also provide an extensive ablation study to validate our different architectural choices and training setup. Finally, we show how the proposed approach can also be naturally extended to integrate observations across different views or body poses, further increasing reconstruction quality. 2. Related work Table 1summarizes the main properties of the most recent generative cloth models we have discussed. Monocular 3D Human Reconstruction is an inherently ill-posed problem and thus greatly benefits from strong human body priors. The reconstructed 3D shape of a human is often a byproduct of 3D pose estimation [16,32,34,35,37, 48,51,59,63,79] represented by a statistical human body end-to-end trainable returns albedo returns shading true surface normals challenging poses animatable semantic editing allows multi-view 3 7 3 3 7 7 7 3 PIFu [61] 7 7 7 3 7 7 7 7 PIFuHD [62] 37377377ARCH++ [25] 7 7 3 3 3 3 7 7 PaMIR [82] 3 3 3 3 7 7 7 7 PHORHUM [6] 7 7 7 3 3 3 7 7 ICON [71] 33333333S3F (Ours) Table 1. Key properties of S3F compared to recent work. Our model includes a number of desirable novel features with respect to previous state-of-the-art, and recovers rigged and relightable human avatars even for challenging body poses in input images. model [40,51,58,73]. Body models, however, only provide mid-resolution body meshes that do not capture important elements of a person’s detail, such as clothing or accessories. To this end, one line of work extends parametric bodies to represent clothing through offsets to the body mesh [1–4,8,49,84]. However, this is prone to fail for loose garments or those with a different topology than the human body. Other representations of the clothed human body have been explored, including voxels [67,83], geometry images [54], bi-planar depth maps [18] or visual hulls [47]. To date, the most powerful representations are implicit functions [6,10,25,26,61,62,71,75,76,82] that define 3D geometry via a decision boundary [13,43] or level set [50]. A popular choice to condition implicit functions on an input image are pixel-aligned features [61]. This approach has been used to obtained detailed 3D reconstruction methods without relying on a body template [6,61,61,75], and in combination with parametric models [25,26,71,75] to take advantage of body priors. ARCH and ARCH++ map pixel-aligned features to a canonical space of SMPL [40] to support animatable reconstructions. However, this requires almost perfect SMPL estimates, which are hard to obtain for images in-the-wild. To address this, ICON [71] proposes an iterative refinement of the SMPL parameters during 3D reconstruction calculations. ICON requires different modules for predicting front & back normal maps, to compute features, and for SMPL fitting. In contrast, we propose an end-to-end trainable model that can correct potential errors in body pose and shape estimation, without supervision, by freely allocating relevant features around the approximate body geometry. Somewhat related are also methods for human relighting in monocular images [28,31,36,65]. However, these methods typically do not reconstruct the 3D human and instead rely on normal maps to transform pixel colors. Our work bears similarity with PHORHUM [6], in our prediction of albedo, and a global scene illumination code. However, we combine this approach with mixed supervision from both synthetic and real data, thus obtaining detailed, photo-realistic reconstructions, for images in-the-wild. Human neural fields. Our work is also related to neural radiance fields, or NeRFs [45], from which we take inspiration for our losses on images in-the-wild. NeRFs have been recently explored for novel human view synthesis [11,15,30,52,53,64,69,70,72]. These methods are trained per subject by minimizing residuals between rendered and observed images and are typically based on a parametric body model. See [66] for an extensive literature review. In contrast, our method is trained on images of different subjects and thus generalizes to unseen identities. Finally, our work follows LVD [16] by predicting a neural deformation field for body vertices, in our case with no explicit supervision. 3. Method We seek to estimate the textured and animation-ready 3D geometry of a person as viewed in a single image. Further, we compute a per-image lighting model to explain shading effects on 3D geometry and to enable relighting for realistic scene placement. Our system utilizes the statistical body model GHUM [73], which represents the human body M(·) as a parametric function of pose θand shape β M(β,θ) : θ×β7→ V∈R3N.(1) GHUM returns a set of 3D body vertices V. We also use imGHUM [5], GHUM’s implicit counterpart. imGHUM computes the signed distance to the body surface sbody xfor any given 3D point xunder a given pose and shape configuration: imGHUM(x,β,θ) : β,θ7→ sbody x. Please refer to the original papers for details. Given a monocular RGB image I(where the person is segmented), together with an approximate 3D geometry represented by GHUM/imGHUM parameters θand β, our method generates an animation-ready, textured 3D reconstruction of that person. We represent the 3D geometry S as the zero-level-set of a signed distance field parameterized by the network f, Sφf=x∈R3|f(I,θ,β,x;φf) = (0,a),(2) with learnable parameters φf. In addition to the signed distance sw.r.t. the 3D surface, fpredicts per-point albedo color a. In the sequel, we denote the signed distance and albedo for point xreturned by fas sxand ax, respectively. Scan be extracted from fby running Marching Cubes [41] in a densely sampled 3D bounding box. In the process of computing sand a,fextracts novel Structured 3D Features. In contrast to pixel-aligned features – a popular representation used by other state-of-the-art methods – our structured 3D features can be re-posed via Linear Blend Skinning (LBS), thus enabling reconstructions in novel poses as well as multi-image feature aggregation. We explain our novel Structured 3D Features in more detail below. 3.1. S3F architecture We divide our end-to-end trainable network into two different parts. First, we introduce our novel Structured 3D Feature (S3F) representation. Second, we describe the procedure to obtain signed distance and color predictions. Finally, we generalize our formulation to multi-view or multiframe settings. Structured 3D Features. Recent methods have shown the suitability of local features for the 3D human reconstruction, due to their ability to capture high-frequency details. Pixel-aligned features [6,26,61,62] are a popular choice, despite having certain disadvantages: (1) it is not straightforward to integrate pixel-aligned features from different time steps or body poses. (2) a ray-depth ambiguity. In order to address these problems, we propose a natural extension of pixel-alignment by lifting the features to the 3D domain: structured 3D features. We obtain the initial 3D feature locations viby densely sampling the surface of the approximate geometry represented by GHUM. We initialize the corresponding features fiby projecting and pooling from a 2D feature map obtained from the input image I g: (I, π(vi)) 7→ fi∈Rk,(3) where π(vi)is a perspective projection of vi, and gis a trainable feature extractor network. The 3D features obtained this way approximate the underlying 3D body but not the actual geometry, especially for loose fitting clothing. To this end, we propose to non-rigidly displace the initial 3D feature locations to better cover the human in the image. Given the initial 3D feature locations vi, we predict per-point displacements in camera coordinates using the network d v0 i=vi+d(fi,ei,vi),(4) where ei∈Rkis a learnt semantic code for vi, initialized randomly and jointly optimized during training. Finally, we project the updated 3D feature locations again to the feature map obtained from I(following the approach from Eq. 3) obtaining our final structured 3D features f0 i. See Fig. 2for an overview. Predicting geometry and texture. We continue describing the process of computing per-point albedo axand signed distances sx. We denote the set of structured 3D features as (F0∈RN×k,V0∈RN×3), where F0refers to the feature vectors and V0to their 3D locations, respectively. First, we use a transformer encoder layer [68] to compute a master feature f? xfor each query point x, using V0as keys and F0 values, respectively. From f? xwe compute per-point albedo and signed distances using a standard MLP t: (x,V0,F0)7→ f? x7→ sx,ax.(5) Illumination code Pooling Structured 3D Features (Eq. 3) Input Image GHUM Estimation Shading AlbedoGeometry Values Keys Query point 3D Volume Image Features Feature extractor Learnable Weights : Per-point learnable code (Eq. 4) (Eq. 5) (Eq. 6) Figure 2. Method overview. We introduce a new implicit 3D representation S3F (Structured 3D Features) that utilizes Npoints sampled on the body surface Vand their 2D projection to pool features Ffrom a 2D feature map extracted from the input image I. The initial body points are non-rigidly displaced by the network dto obtain V0in order to sample new features F0, on locations that are not covered by body vertices, such as loose clothing or hair. Given an input point x, we then aggregate representations from the set of points and their features using a transformer architecture t, to finally obtain per-point signed distance and albedo color. Finally, we can relight the reconstruction using the predicted albedo, an illumination representation Land a shading network p. We assume perspective projection to obtain more natural reconstructions, with correct proportions. While previous work maps query points xto a continuous feature map (e.g. pixel-aligned features), we instead aim to pool features from the discrete set F0. Intuitively, the most informative features for xshould be located in its 3D neighborhood. We therefore first explore a simple baseline, by collecting the features of the three closest points of x in V0and interpolating them using barycentric coordinates. However, we found this approach not sufficient (see Tab. 2). With keys and queries based on 3D positions, we argue that a transformer encoder is a useful tool to integrate relevant structured 3D features based on learned distance-based attention. In practice, we use two different transformer heads and MLPs to predict albedo and distance, please see Sup. Mat. for details. Furthermore, we compute signed distances sxby updating the initial estimate of imGHUM with a residual, sx=sbody x+ ∆sx. This in practice makes the training more stable for challenging body poses. To model scene illumination, we follow PHORHUM [6] and utilize the bottleneck of the feature extractor network gas scene illumination code L. We then predict per-point shading coefficient δxusing a surface shading network p p(nx,L;φp)7→ δx,(6) parameterized by weights φp, where the normal of the input point nx=∇xsxis the estimated surface normal defined by the gradient of the estimated distance w.r.t. x. The final shaded color is then obtained as cx=δxaxwith  denoting element-wise multiplication. 3D Feature Manipulation & Aggregation. Since the proposed approach relies on 3D feature locations originally sampled from GHUM’s body surface, we can utilize GHUM to re-pose the representation. This has two main benefits: (1) we can reconstruct in a pose different from the one in the image, e.g. by resolving self-contact, which is typically not possible after creating a mesh, and (2) we can aggregate information from several observations as follows. In a multi-view or multi-frame setting with Oobservations of the same person under different views or poses, we define each input image as It,∀t∈ {1, . . . , O}. Let us extend the previous notation to (V0 t,F0 t)as structured 3D features for a given frame t. To integrate features from all images, we use LBS to invert the posing transformation and map them to GHUM’s canonical pose denoted as ˜ V0. For each observation, we compute point visibilities Ot∈RN×1(based on the GHUM mesh) and weight all feature vectors and canonical feature positions by the overall normalized visibility, thus obtaining an aggregated set of structured features ˆ V0= T X t softmaxt(Ot)˜ V0 t,(7) ˆ F0= T X t softmaxt(Ot)˜ F0 t(8) where softmax normalizes per-point contributions along all views. The aggregated structured 3D features (ˆ V0,ˆ F0) can be posed to the original body pose of each input image, or to any new target pose. Finally, we run the remaining part of the model in the posed space to predict sxand ax. 3.2. Training S3F The semi-supervised training pipeline leverages a small subset of 3D synthetic scans together with images in-the- wild. During training, we run a forward pass for both inputs and integrate the gradients to run a single backward step. Real data. We use images in-the-wild paired with body pose and shape parameters obtained via fitting. We su- Geometry Color (PSNR) Method Chamfer ↓IoU ↑NC↑Albedo ↑Shading ↑ Architecture: Using pixel-aligned features 3.59 0.360 0.828 13.17 14.13 + Predicting SDF residual 5.43 0.428 0.851 13.19 13.06 + Lifting features to 3D 0.606 0.644 0.765 9.82 10.07 + Transformer 0.532 0.720 0.911 16.54 15.85 FULL: + Point displacement 0.339 0.734 0.924 16.31 16.67 Supervision regime: FULL: Only synthetic data 0.381 0.719 0.916 16.34 15.84 FULL: Only real data 0.444 0.723 0.907 10.96 13.59 Table 2. Ablation of several of our design choices. Chamfer metrics are ×10−3. pervise on the input images through color and occupancy losses. Following recent work on neural rendering of combined signed distance and radiance fields [77], we convert predicted signed distance values sxinto density such that σ(x) = β−1Ψβ(−sx),(9) where Ψβ(·)is the CDF of the Laplace distribution, and βis a learnable scalar parameter equivalent to a sharpness factor. After training, we no longer need βfor reconstruction but can use it to render novel views (see Sec. 4.3). The color of a specific image pixel is estimated via the volume rendering integral by accumulating shaded colors and volume densities along its corresponding camera ray (see [45] or Sup. Mat. for more details). After rendering, we minimize the residual between the ground-truth pixel color crand the accumulation of shaded colors ˆ cralong the ray of pixel rsuch that Lrgb =kcr−ˆ crk1. Additionally, we also define a VGG-loss [12]Lvgg over randomly sampled front patches, enforcing the structure to be more realistic. In addition to color, we also supervise geometry by minimizing the difference between the integrated density ˆσrand the ground truth pixel mask σr:Lmask =kσr−ˆσrk1. Finally, we regularize the geometry using the Eikonal loss [22] on a set of points Ωsampled near the body surface such that Leik =X x∈Ω (k∇xsxk2−1)2.(10) The full loss for the real data is a linear combination of the previous components with weights λ∗Lreal =Lrgb + λvggLvgg +λmaskLmask +λeikLeik. (Quasi-)Synthetic data. Most previous work relies only on quasi-synthetic data in the form of re-rendered textured 3D human scans, for 3D human reconstruction. While we aim to alleviate the need for synthetic data, it is useful for supervising features not observable in images in-the-wild, such as albedo, color, or geometry of unseen body parts. To this end, we additionally supervise our model on a small set of pairs of 3D scans and synthetic images. For the quasisynthetic subset, we use the same losses as for the real data and add additional supervision from the 3D scans. Given a Geometry Color (PSNR) Method Chamfer ↓IoU ↑NC ↑Albedo ↑Shading ↑ GHUM [73]3.56 0.562 0.750 - - PIFu [61]6.61 0.519 0.738 - 11.15 Geo-PIFu [24]9.61 0.453 0.710 - 10.62 PIFuHD [62]4.94 0.552 0.749 - - PaMIR [82]5.35 0.597 0.763 - - ARCH [26]7.52 0.549 0.712 - 10.66 ARCH++ [25]6.53 0.549 0.722 - 10.37 ICON [71]3.53 0.622 0.785 - - ICON [71]+[44]4.47 0.599 0.764 - - PHORHUM [6]2.92 0.594 0.814 11.97 11.20 Ours (No shading) 2.27 0.675 0.827 - 16.62 Ours 1.88 0.694 0.847 15.06 14.81 Ours (GT pose/shape) 0.339 0.734 0.924 16.31 16.67 Table 3. Quantitative comparison against other monocular 3D human reconstruction methods. Chamfer metrics are ×10−3 and PSNR is obtained from all 3D scan vertices, including those not visible in the input image. 3D scan, we sample Bpoints on the surface together with its ground truth albedo aGT iand shaded color sGT i. We minimize the difference in predicted albedo and shaded color such that L3D rgb =Px∈BkaGT x−ˆ axk1+kcGT x−ˆ cxk1. We additionally sample a set of 3D points Ωclose to the scan surface and compute inside/outside labels lfor our final loss, L3D label =Px∈ΩBCE(lx, σ(x)), where BCE is the Binary Cross Entropy function. The overall loss for synthetic data Lsynth is a linear combination of all losses described above. We train our model by minimizing both real and synthetic objectives Ltotal =Lreal +Lsynth. Implementation details are available in the Sup. Mat. 4. Experiments Our goal is to design a single 3D representation and an end-to-end semi-supervised training methodology that is flexible enough for a number of tasks. In this section, we evaluate our model’s performance on the tasks of monocular 3D human reconstruction, albedo, and shading estimation, as well as for novel view rendering and reconstruction from both single-images and video. Finally, we illustrate our model’s capabilities for 3D garment editing. 4.1. Data We follow a mixed supervision approach by combining synthetic data with limited clothing or pose variability, with in-the-wild images of people in challenging poses/scenes. Synthetic data. We use 35 rigged and 45 posed scans from RenderPeople [57] for training. We re-pose the rigged scans to 200 different poses and render them five times with different HDRI backgrounds and lighting using Blender [9]. With probability 0.4we use a frontal view, otherwise we render the person from a random azimuth and a uniform random elevation in [−20,20]◦. We obtain ground-truth albedo directly from the scan’s texture map, and shade the C Input image ICON [71] OURS Input image ICON [71] OURS Figure 3. Qualitative 3D human reconstruction comparisons against the SOTA method ICON [71]for complex body poses. Images on the right column feature challenging poses and viewpoints. Notice the additional detail provided by our method, especially for hair and faces. We additionally recover per-vertex albedo (not shown) and shaded color. See Sup. Mat. for more examples and failure cases. full scan (including occluded regions) to obtain groundtruth shaded colors before rendering. Images in-the-wild. We use the HITI dataset [7], a collection of images in-the-wild with 2D key-point annotations and human segmentation masks, to which we fit GHUM by minimizing the 2D joint reprojection error. We automatically remove images with high fitting error, resulting in 40K images for training. We use the provided test set to validate results. More details are provided in the Sup. Mat. Evaluation data. We compare our method for monocular reconstruction using a synthetically generated test set based on a test-split of 3D scans, featuring significant diversity of body poses (e.g. running), illumination, and viewpoints. For images in-the-wild, we obtain GHUM parameter initialisation using [21] and further optimize by minimizing 2D joint reprojection errors. For the task of novel-view rendering we use GHS3D [73], which consists of 14 multi-view videos of different people, between 30-60 frames each, with 4 views for training and 2 for testing. We evaluate on test views and compare results based on either a single training image or full video. Note that our method does not ”train”, but only computes a forward pass based on a training image. 4.2. Monocular 3D Human Reconstruction Ablation Study. We ablate our main methodological choices in Table 2. We use ground-truth pose and shape parameters for the test set, to factor out potential errors in body estimation that otherwise dominate the reconstruction metrics. We provide different baselines to demonstrate the contribution at each step. Albedo and shading evaluation is based on nearest neighbors of the scan, including the back or non-visible regions, which tend to dominate PSNR error rates. In the first two rows we show results of pixelaligned only baselines. By lifting features to the 3D domain and placing them around the body, our method already achieves competitive performance. This is in line with other works [25,26,71] which showed that extending a statistical body model leads to better reconstruction for challenging poses and viewpoints. We improve on this by using a transformer architecture associating query points with 3D features, especially when using our S3F (full model). Finally, we compare our performance when training only on real or Image Initial points S3Fs Geometry Albedo Shaded Figure 4. Output of S3F. Given an input image (first col.) and GHUM pose estimates (second col.), S3F displaces points efficiently to cover relevant features from the image (third col.). Displaced points cover areas that not necessarily belong to the foreground, but tend to represent loose clothing (e.g. hair and loose clothing in first and second examples respectively). This coverage and the efficient feature retrieval allows to correct potential errors in the body such as the feet in the first two examples or the overall body orientation in the third row. Zoom-in recommended. Input Image Albedo Shaded Albedo Shaded PHORHUM [6] OURS Figure 5. Qualitative comparison against the state-of-the-art method PHORHUM [6]for albedo and shading estimation from monocular images. synthetic data and show that we get the best results when using a mix of both. Moreover, real data helps generalization to in-the-wild images, not captured in our test set. SOTA Comparison. We evaluate our method for monoc- Chamfer↓IoU↑NC↑PSNR↑SSIM↑LPIPS↓Train time NeuralBody [53]0.790 0.887 0.810 24.70 0.829 0.236 hours H-NeRF [72]0.218 0.932 0.890 24.92 0.852 0.232 hours Ours (finetuned) 0.244 0.841 0.910 27.54 0.871 0.148 30 min One-shot novel view rendering Ours (1 image) 0.504 0.795 0.899 23.75 0.825 0.167 0 Ours (video) 0.473 0.807 0.905 24.82 0.836 0.148 0 Table 4. Quantitative comparison against recent novel view rendering methods for the GHS3D Dataset [72]. We achieve competitive reconstruction, and view rendering metrics using just one forward pass, especially when taking multiple frames as input. This validates our proposed feature integration scheme. ular 3D human reconstruction and compare with previous methods quantitatively in Table 3. Here, we do not assume ground-truth pose/shape for each image, and instead obtain them from the input images alone, for all methods. We observe that while metrics are dominated by errors in pose estimation, our method is more robust and can correct these errors by using S3Fs. Additionally, our model recovers albedo and shading significantly better than baselines. Qualitatively, we compare against ICON [71] in Fig. 3, for different levels of body pose complexity. Our method is more robust to both challenging poses and loose clothing than previous SOTA while additionally predicting color, being animatable and relightable. We provide additional comparisons to other methods in Sup. Mat. We also compare our method qualitatively against PHORHUM [6] on the task of albedo and shading estimation from monocular images in Fig. 5. We observe that our method is more robust and we hypothesize this is due to PHORHUM’s trained on synthetic data only, which impacts its ability to generalize to images in-the-wild. We show more results of our method in Fig. 4, where we also visualize the initial and displaced 3D locations of our body points. While the initial GHUM fit might lead to noisy estimates, the displacement step is able to correct errors (feet in first & second rows) and better cover loose clothing (second row). 4.3. Applications The main applications of human digitization (e.g. VR/AR or telepresence) require considerable control over virtual avatars, in order to animate, relight and edit them. Here, we evaluate the performance of our approach for novel view rendering to a full video, given as little as one input image. This process entails the ability to re-pose reconstructions and integrate features from different views or video frames. We then showcase the model’s potential for relighting and clothing editing. Novel view synthesis. For novel-view synthesis, previous work often requires tens of GPU hours to train personspecific networks. Here, we compare against two such recent works, NeuralBody [53] and H-Nerf [72], on the GHS3D dataset [72]. We report reconstruction and novel Groundtruth Novel view & pose From 1 view From video Finetuned Input image GT view & pose Input Image From 1 view From video Finetuned Figure 6. 3D Human reconstruction from one image, video and after finetuning our network geometry and color heads for 30 minutes. The image used as input in the single-view case is shown in the second column. By using only this information, previously occluded body areas might lead to larger errors when rendering from novel viewpoints or in a different pose, e.g. first example, where face and frontal details need to be hallucinated. This can be corrected by integrating information from multiple frames. Input image Relighted Reconstructions Figure 7. Qualitative results on 3D Human Relighting. view metrics in Table 4. These show that our method obtains comparable performance without retraining, just from a forward pass. For video, we have enough supervision to robustly fine-tune the geometric head of our network, in just 30 minutes, to obtain a significant performance boost in both geometry and rendering metrics. In Fig. 6we show the difference between reconstructing from a single view or full video. Even though our method yields a reasonable 3D reconstruction from just a single image, it has to hallucinate color in occluded areas (e.g. face and front body parts in the first row). The resulting inconsistency can be largely resolved when novel views are considered. Relighting. By predicting per-point albedo and a scene illumination, our model returns a relightable reconstruction of the person in the target image. We show a number of relighted reconstructions in Fig. 7given a target input im- Image Reconstruction Cloth texture transfer Figure 8. Clothing texture transfer. Given an image of a target person, we identify S3Fs that project inside upper-body cloth segmentation [38], and replace their feature vectors for those obtained from other subjects. More examples, cloth types and reference images and segmentation masks are shown in Sup. Mat. age (left column). The reconstructions are consistently reshaded in novel scenes, thus enabling 3D human compositing applications. Clothing editing. We harvest the semantic properties of the proposed Structured 3D Features to explore potential applications for 3D virtual try-on. Given an input image of a person and a segmentation mask [38] of a particular piece of clothing (e.g. upper-body), we first find body points that project inside the target mask. By exchanging their features with those obtained from other images, we can effectively transfer clothing texture. Fig. 8showcases diverse clothing editing cases for the upper-body. See Sup. Mat. for other subjects and garments. While 3D clothing transfer is a complex problem by itself [46,74], we tackle a much more challenging scenario than previous methods that typically map color to a known 3D template; and we use a single, versatile end-to-end model for monocular 3D reconstruction, relighting, and 3D editing. 5. Discussion Limitations. Failures in monocular human reconstruction occur mainly due to incorrect GHUM fits or very loose clothing. The lack of ambient occlusions in the proposed scheme might lead to incorrect albedo estimates for images in-the-wild. Misalignment in estimated GHUM pose & shape might lead to blurriness, especially for faces. We show failure cases in Supp. Mat. Ethical considerations. We present a human digitization tool that allows relighting and re-posing avatars. However, neither is it intended, nor particularly useful for any form of deep fakes, since the quality of reconstructions is not at par with facial deep fakes. We aim to improve model coverage for diverse subject distributions by means of a weakly supervised approach and by taking advantage of extensive image collections, for which labeled 3D data may be difficult or impossible to collect. Conclusion. We have presented a controllable transformer methodology, S3F, relying on versatile attention-based features, that enables the 3D reconstruction of rigged and relightable avatars by means of a single semi-supervised end- to-end trainable model. Our experimental results illustrate how S3F models obtain state-of-the-art performance for the task of 3D human reconstruction and for albedo and shading estimation. We also show the potential of our models for 3D virtual try-on applications. Supplementary Material In this supplementary, we describe the implementation details of our method extensively and provide more results and failure cases. We also include a Supplementary Video summarizing our contributions and results. A. Implementation Details Data. Our synthetic data is based on a set of 3D Render- People Scans [57]. We use 35 rigged and 45 posed scans for training. The rigged scans are re-posed to 200 different poses sampled randomly from the CMU motion sequences [14]. With probability 0.4we render a scan from a frontal view, otherwise, we render from a random azimuth and a uniform random elevation in [−20,20]◦. We render each scan using high dynamic range image (HDRI) [23] lighting and backgrounds using Blender [9]. We obtaining ground-truth albedo directly from the scan’s texture map and bake the full scan’s shading (including occluded regions) to obtain ground-truth shaded colors. For real images, we use the HITI dataset [7], which contains 150K images in-the-wild with predicted foreground segmentation masks and annotated 2D human keypoints. We obtained initial GHUM parameters for each image by estimating pose and shape using [21]. We then further optimized pose and shape parameters by minimizing the 2D reprojection error, the normalizing flow pose prior from [73,78] and a body shape prior. The weights for joint reprojection, body pose regularization and body shape regularization are 10,1and 10 respectively, and we assumed a perspective camera projection with fixed focal length. After fitting, we remove fits that have an average reprojection error greater than 3 pixels or where at least 10% of the body surface projects outside the segmentation mask, leading to 40k training images. For inference on images-in-the-wild we follow the same fitting procedure, by leveraging predicted 2D keypoints instead of ground-truth annotations. Architecture. We take masked input images at 512×512 px resolution. We augment the input with rendered normal and semantic maps to provide information about the GHUM fit to the feature extractor network. The semantic map is obtained by rendering the original template vertex locations as vertex colors, essentially defining a dense correspondence map to GHUM’s zero-pose. We normalize both maps between 0 and 1 and stack them with the original image before passing all to the image feature extractor network. We noticed that by concatenating normal and semantic maps, the network is better able to correct noisy GHUM fits and their geometry. The feature extractor network is a U-Net [60] with 6 encoder and 7 decoder layers with sizes [64,128,256,512,512,512] and [512,512,512,512,256,256,256] respectively. The illumination code is extracted from the bottleneck and has shape 8×8×512. The output per-pixel feature maps is of shape 512 ×512 ×256, with 256-feature vectors. We next detail how we sample points on the body surface and pool from image features. We explored different point densities sampled from the body surface, all based on the original GHUM mesh to maintain correspondences. To subdivide a mesh, we add a vertex in the center of each edge, increasing the number of faces by a factor of 4. The subdivision is fast to perform at train/test time and does not cause any significant overhead. However, naively subdividing points on the mesh leads to a large number of body points causing memory issues in later stages (e.g. processing them on the transformer encoder), thus we additionally run K-Means on the subdivided template mesh obtaining clusters of 2k, 5k, 8k, 10k, 12k, 15k and 18k body points. We ablate qualitatively the number of points on the model in Fig. 19. Using fewer points does not significantly affect the results, but produces blurrier color reconstructions. Our final model uses 18k points. During inference, given the GHUM fit of an image and estimated camera parameters, we project these points to the image and extract image features. We use the last 64 features of the previously extracted image features to predict per-vertex deformation, using a 2- Layer MLP with hidden shapes [64,3] and leaky-ReLU after the first layer. We project the deformed vertices again to obtain the remaining 192 per-pixel features. The goal of the transformer encoder is to efficiently map a query point xto the Structured 3D Features. We first map both deformed body points and query point positions to a higher-dimensional space using positional encoding with 6 frequencies, and apply a shared 2-Layer MLP with output size 256. In practice, we use two MLPs predicting geometry and albedo independently. We next apply attention [68] to combine per-point features and obtain f? x. The final geometry and color heads are both MLPs with eight 512-dimensional fully-connected layers and Swish activation [55], an output layer with Sigmoid activation for the color component, and a skip connection to the fourth layer. The shading network sis conditioned on the previously extracted illumination code and consists of three 256- dimensional fully-connected layers with ReLU activation, including the output layer. The weights of all layers are initialized with Xavier initialization [19], with the βparameter (Eq. 9) being initialized as 0.1. We train all network components jointly end- to-end for 500k iterations using the Adam optimizer [33], with an initial learning-rate of 1×10−4that linearly de- Input image Relighted Reconstructions Figure 15. Qualitative results on 3D Human Relighting. Input image Animated Reconstructions Figure 16. Qualitative results on animation of 3D reconstructions. Source clothing Input image Figure 17. More examples of cloth texture transfer. We extend the experiment from the main paper, showcasing an example of upperbody cloth try-on, and show the input images (left) and source clothing (upper-row). The 3D reconstructions look realistic and consistent accross all examples. Note that we only take one single image from both subject and clothing, and the network re-poses the affected S3Fs, allucinates occluded texture, and shades the clothing in the new scene. Source clothing Input image Figure 18. More examples of cloth texture transfer. This example features try-on examples from lower-body clothing. See the Supplementary Video for more examples. Input image 2000 5000 8000 10000 12000 15000 18000 Number of sampled points in the body surface Figure 19. Qualitative ablation of number of points sampled in the body surface, storing 3D Features. Our final model has 18000 points which provides the best tradeoff between GPU memory and sharpness. Lower number of points do not affect significantly quantitative results but are less capable of representing high-frequency details. References [1] Thiemo Alldieck, Marcus Magnor, Bharat Lal Bhatnagar, Christian Theobalt, and Gerard Pons-Moll. Learning to reconstruct people in clothing from a single rgb camera. In CVPR, 2019. 1,2 [2] Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Detailed human avatars from monocular video. In 3DV, 2018. 2 [3] Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruction of 3d people models. In CVPR, 2018. 2 [4] Thiemo Alldieck, Gerard Pons-Moll, Christian Theobalt, and Marcus Magnor. Tex2shape: Detailed full human body geometry from a single image. In ICCV, 2019. 1,2 [5] Thiemo Alldieck, Hongyi Xu, and Cristian Sminchisescu. imGHUM: Implicit generative models of 3D human shape and articulated pose. In ICCV, 2021. 3 [6] Thiemo Alldieck, Mihai Zanfir, and Cristian Sminchisescu. Photorealistic monocular 3d reconstruction of humans wearing clothing. In CVPR, 2022. 1,2,3,4,5,7,10 [7] Eduard Gabriel Bazavan, Andrei Zanfir, Mihai Zanfir, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. Hspace: Synthetic parametric humans animated in complex environments. arXiv, 2021. 6,9,11 [8] Bharat Lal Bhatnagar, Garvita Tiwari, Christian Theobalt, and Gerard Pons-Moll. Multi-garment net: Learning to dress 3d people from images. In ICCV, 2019. 2 [9] Blender Online Community. Blender - a 3D modelling and rendering package. Blender Foundation, Blender Institute, Amsterdam, 2020. 5,9 [10] Yukang Cao, Guanying Chen, Kai Han, Wenqi Yang, and Kwan-Yee K. Wong. Jiff: Jointly-aligned implicit face function for high quality single view clothed human reconstruction. In CVPR, 2022. 2 [11] Jianchuan Chen, Ying Zhang, Di Kang, Xuefei Zhe, Linchao Bao, Xu Jia, and Huchuan Lu. Animatable neural radiance fields from monocular rgb videos, 2021. 3 [12] Qifeng Chen and Vladlen Koltun. Photographic image synthesis with cascaded refinement networks. In ICCV, 2017. 5, 10 [13] Zhiqin Chen and Hao Zhang. Learning implicit fields for generative shape modeling. In CVPR, 2019. 2 [14] Cmu graphics lab motion capture database. http:// mocap.cs.cmu.edu/.9 [15] Enric Corona, Tomas Hodan, Minh Vo, Francesc Moreno- Noguer, Chris Sweeney, Richard Newcombe, and Lingni Ma. Lisa: Learning implicit shape and appearance of hands. CVPR, 2022. 3 [16] Enric Corona, Gerard Pons-Moll, Guillem Aleny` a, and Francesc Moreno-Noguer. Learned vertex descent: A new direction for 3d human model fitting. ECCV, 2022. 2,3 [17] Enric Corona, Albert Pumarola, Guillem Alenya, Gerard Pons-Moll, and Francesc Moreno-Noguer. Smplicit: Topology-aware generative model for clothed people. In CVPR, 2021. 1 [18] Valentin Gabeur, Jean-S´ ebastien Franco, Xavier Martin, Cordelia Schmid, and Gregory Rogez. Moulding humans: Non-parametric 3d human shape estimation from single images. In ICCV, 2019. 2 [19] Xavier Glorot and Yoshua Bengio. Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics, pages 249–256. JMLR Workshop and Conference Proceedings, 2010. 9 [20] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial networks. Communications of the ACM, 63(11):139–144, 2020. 10 [21] Ivan Grishchenko, Valentin Bazarevsky, Andrei Zanfir, Eduard Gabriel Bazavan, Mihai Zanfir, Richard Yee, Karthik Raveendran, Matsvei Zhdanovich, Matthias Grundmann, and Cristian Sminchisescu. Blazepose ghum holistic: Realtime 3d human landmarks and pose estimation. arXiv, 2022. 6,9 [22] Amos Gropp, Lior Yariv, Niv Haim, Matan Atzmon, and Yaron Lipman. Implicit geometric regularization for learning shapes. ICML, 2020. 5,10 [23] https://polyhaven.com/.9 [24] Tong He, John Collomosse, Hailin Jin, and Stefano Soatto. Geo-pifu: Geometry and pixel aligned implicit functions for single-view human reconstruction. NeurIPS, 2020. 5 [25] Tong He, Yuanlu Xu, Shunsuke Saito, Stefano Soatto, and Tony Tung. Arch++: Animation-ready clothed human reconstruction revisited. In CVPR, 2021. 1,2,5,6 [26] Zeng Huang, Yuanlu Xu, Christoph Lassner, Hao Li, and Tony Tung. Arch: Animatable reconstruction of clothed humans. In CVPR, 2020. 1,2,3,5,6 [27] Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. Perceiver: General perception with iterative attention. In International conference on machine learning, pages 4651–4664. PMLR, 2021. 10 [28] Chaonan Ji, Tao Yu, Kaiwen Guo, Jingxin Liu, and Yebin Liu. Geometry-aware single-image full-body human relighting. In ECCV, 2022. 2 [29] Boyi Jiang, Yang Hong, Hujun Bao, and Juyong Zhang. Selfrecon: Self reconstruction your digital avatar from monocular video. In CVPR, 2022. 1 [30] Wei Jiang, Kwang Moo Yi, Golnoosh Samei, Oncel Tuzel, and Anurag Ranjan. Neuman: Neural human radiance field from a single video. ECCV, 2022. 3 [31] Yoshihiro Kanamori and Yuki Endo. Relighting humans: occlusion-aware inverse rendering for full-body human images. SIGGRAPH Asia, 2019. 2 [32] Angjoo Kanazawa, Michael J Black, David W Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In CVPR, 2018. 2 [33] Diederik P Kingma and Jimmy Ba. Adam: A method for stochastic optimization. arXiv, 2014. 9 [34] Nikos Kolotouros, Georgios Pavlakos, Michael J Black, and Kostas Daniilidis. Learning to reconstruct 3d human pose and shape via model-fitting in the loop. In ICCV, 2019. 2 [35] Nikos Kolotouros, Georgios Pavlakos, and Kostas Daniilidis. Convolutional mesh regression for single-image human shape reconstruction. In CVPR, 2019. 2 [36] Manuel Lagunas, Xin Sun, Jimei Yang, Ruben Villegas, Jianming Zhang, Zhixin Shu, Belen Masia, and Diego Gutierrez. Single-image full-body human relighting. ECCV, 2022. 2 [37] Christoph Lassner, Javier Romero, Martin Kiefel, Federica Bogo, Michael J Black, and Peter V Gehler. Unite the people: Closing the loop between 3d and 2d human representations. In CVPR, 2017. 2 [38] Peike Li, Yunqiu Xu, Yunchao Wei, and Yi Yang. Selfcorrection for human parsing. PAMI, 2020. 8 [39] Tianye Li, Shichen Liu, Timo Bolkart, Jiayi Liu, Hao Li, and Yajie Zhao. Topologically consistent multi-view face inference using volumetric sampling. In ICCV, 2021. 1 [40] Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. Smpl: A skinned multiperson linear model. ToG, 2015. 2 [41] William E Lorensen and Harvey E Cline. Marching cubes: A high resolution 3d surface construction algorithm. SIGGRAPH, 1987. 3,10 [42] Lars Mescheder, Andreas Geiger, and Sebastian Nowozin. Which training methods for gans do actually converge? In International conference on machine learning, pages 3481– 3490. PMLR, 2018. 10 [43] Lars Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. Occupancy networks: Learning 3d reconstruction in function space. In CVPR, 2019. 2 [44] Marko Mihajlovic, Aayush Bansal, Michael Zollhoefer, Siyu Tang, and Shunsuke Saito. KeypointNeRF: Generalizing image-based volumetric avatars using relative spatial encoding of keypoints. In ECCV, 2022. 5 [45] Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view synthesis. ECCV, 2020. 3,5 [46] Aymen Mir, Thiemo Alldieck, and Gerard Pons-Moll. Learning to transfer texture from clothing images to 3d humans. In CVPR, 2020. 8 [47] Ryota Natsume, Shunsuke Saito, Zeng Huang, Weikai Chen, Chongyang Ma, Hao Li, and Shigeo Morishima. Siclope: Silhouette-based clothed people. In CVPR, 2019. 2 [48] Mohamed Omran, Christoph Lassner, Gerard Pons-Moll, Peter Gehler, and Bernt Schiele. Neural body fitting: Unifying deep learning and model based human pose and shape estimation. In 3DV, 2018. 2 [49] Hayato Onizuka, Zehra Hayirci, Diego Thomas, Akihiro Sugimoto, Hideaki Uchiyama, and Rin-ichiro Taniguchi. Tetratsdf: 3d human reconstruction from a single image with a tetrahedral outer shell. In CVPR, 2020. 2 [50] Jeong Joon Park, Peter Florence, Julian Straub, Richard Newcombe, and Steven Lovegrove. Deepsdf: Learning continuous signed distance functions for shape representation. In CVPR, 2019. 2 [51] Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed AA Osman, Dimitrios Tzionas, and Michael J Black. Expressive body capture: 3d hands, face, and body from a single image. In CVPR, 2019. 2 [52] Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Animatable neural radiance fields for modeling dynamic human bodies. In CVPR, 2021. 3 [53] Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR, 2021. 1,3,7 [54] Albert Pumarola, Jordi Sanchez-Riera, Gary Choi, Alberto Sanfeliu, and Francesc Moreno-Noguer. 3dpeople: Modeling the geometry of dressed humans. In ICCV, 2019. 2 [55] Prajit Ramachandran, Barret Zoph, and Quoc V Le. Searching for activation functions. arXiv preprint arXiv:1710.05941, 2017. 9 [56] Edoardo Remelli, Timur Bagautdinov, Shunsuke Saito, Chenglei Wu, Tomas Simon, Shih-En Wei, Kaiwen Guo, Zhe Cao, Fabian Prada, Jason Saragih, et al. Drivable volumetric avatars using texel-aligned features. In SIGGRAPH, 2022. 1 [57] Renderpeople dataset. https : / / renderpeople . com/.5,9 [58] Javier Romero, Dimitrios Tzionas, and Michael J Black. Embodied hands: Modeling and capturing hands and bodies together. ToG, 2017. 2 [59] Yu Rong, Takaaki Shiratori, and Hanbyul Joo. Frankmocap: Fast monocular 3d hand and body motion capture by regression and integration. arXiv, 2020. 2 [60] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U- net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, pages 234–241. Springer, 2015. 9 [61] Shunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima, Angjoo Kanazawa, and Hao Li. Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. In ICCV, 2019. 1,2,3,5,10 [62] Shunsuke Saito, Tomas Simon, Jason Saragih, and Hanbyul Joo. Pifuhd: Multi-level pixel-aligned implicit function for high-resolution 3d human digitization. In CVPR, 2020. 1,2, 3,5,10 [63] David Smith, Matthew Loper, Xiaochen Hu, Paris Mavroidis, and Javier Romero. Facsimile: Fast and accurate scans from an image in less than a second. In ICCV, 2019. 2 [64] Shih-Yang Su, Frank Yu, Michael Zollhoefer, and Helge Rhodin. A-nerf: Surface-free human 3d pose refinement via neural rendering. NeurIPS, 2021. 3 [65] Daichi Tajima, Yoshihiro Kanamori, and Yuki Endo. Relighting humans in the wild: Monocular full-body human relighting with domain adaptation. In CGF, 2021. 2 [66] Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, W Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, et al. Advances in neural rendering. In CGF, 2022. 3 [67] Gul Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid. Bodynet: Volumetric inference of 3d human body shapes. In ECCV, 2018. 2 [68] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. NeurIPS, 2017. 3,9 [69] Shaofei Wang, Katja Schwarz, Andreas Geiger, and Siyu Tang. Arah: Animatable volume rendering of articulated human sdfs. In ECCV, 2022. 3 [70] Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Humannerf: Free-viewpoint rendering of moving people from monocular video. In CVPR, 2022. 3 [71] Yuliang Xiu, Jinlong Yang, Dimitrios Tzionas, and Michael J Black. Icon: Implicit clothed humans obtained from normals. In CVPR, 2022. 1,2,5,6,7 [72] Hongyi Xu, Thiemo Alldieck, and Cristian Sminchisescu. H-nerf: Neural radiance fields for rendering and temporal reconstruction of humans in motion. NeurIPS, 2021. 1,3,7 [73] Hongyi Xu, Eduard Gabriel Bazavan, Andrei Zanfir, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. Ghum & ghuml: Generative 3d human shape and articulated pose models. In CVPR, 2020. 2,3,5,6,9 [74] Xiangyu Xu and Chen Change Loy. 3d human texture estimation from a single image with transformers. In CVPR, 2021. 8 [75] Ze Yang, Shenlong Wang, Sivabalan Manivasagam, Zeng Huang, Wei-Chiu Ma, Xinchen Yan, Ersin Yumer, and Raquel Urtasun. S3: Neural shape, skeleton, and skinning fields for 3d human modeling. In CVPR, 2021. 2 [76] Kennard Yanting Chan, Guosheng Lin, Haiyu Zhao, and Weisi Lin. Integratedpifu: Integrated pixel aligned implicit function for single-view human reconstruction. In ECCV, 2022. 2 [77] Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. Volume rendering of neural implicit surfaces. NeurIPS, 2021. 5 [78] Andrei Zanfir, Eduard Gabriel Bazavan, Hongyi Xu, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. Weakly supervised 3d human pose and shape reconstruction with normalizing flows. In ECCV, pages 465– 481. Springer, 2020. 9 [79] Mihai Zanfir, Andrei Zanfir, Eduard Gabriel Bazavan, William T Freeman, Rahul Sukthankar, and Cristian Sminchisescu. Thundr: Transformer-based 3d human reconstruction with markers. In CVPR, 2021. 2 [80] Jingyang Zhang, Yao Yao, Shiwei Li, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. Critical regularizations for neural surface reconstruction in the wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6270–6279, 2022. 10 [81] Yufeng Zheng, Victoria Fern´ andez Abrevaya, Marcel C B¨ uhler, Xu Chen, Michael J Black, and Otmar Hilliges. Im avatar: Implicit morphable head avatars from videos. In CVPR, 2022. 1 [82] Zerong Zheng, Tao Yu, Yebin Liu, and Qionghai Dai. Pamir: Parametric model-conditioned implicit representation for image-based human reconstruction. PAMI, 2021. 1, 2,5 [83] Zerong Zheng, Tao Yu, Yixuan Wei, Qionghai Dai, and Yebin Liu. Deephuman: 3d human reconstruction from a single image. In ICCV, 2019. 2 [84] Hao Zhu, Xinxin Zuo, Sen Wang, Xun Cao, and Ruigang Yang. Detailed human shape estimation from a single image by hierarchical mesh deformation. In CVPR, 2019. 2