Human-Hair Priors Rebuild 3D Animal Fur 10x Faster, No Animal-Fur Dataset Needed

FurE: Efficient Instance-Specific 3D Fur Reconstruction without Animal-Fur Datasets

Srinjay Sarkar, Prakhar Kaushik, Soumava Paul, Alan Yuille

cs.CV, cs.AI, cs.GR

2026-09-29

FurE rebuilds editable strand-level animal fur from multi-view images using a human-hair PCA prior, cutting strand training from 10.5 hours to 52 minutes and surviving a real bison sequence where NeuralFur fails.

What problem this solves

Production pipelines represent hair and fur as explicit strands, individual curves that artists can edit, render, and simulate. Reconstructing that representation from photos is established for human hair, largely because human-hair datasets exist to train priors. Animals have no equivalent dataset, fur covers most of the body, and it varies across species and across body parts of the same animal. Dense self-occlusion means multi-view reconstruction only recovers the outer furry envelope, never the skin where strand roots should attach.

The closest prior work, NeuralFur (3DV 2026), fills these gaps with SMAL template fitting and VLM-derived fur attributes, but optimizes every strand densely: 10.5 hours of strand training plus 10 hours of preprocessing per instance. At that cost, per-instance animal capture stays a lab demo.

Method

FurE runs in two stages, and the organizing idea is to move the expensive part of the problem from full 3D strand space into a low-dimensional latent space.

Stage one finds the skin. After NeuS2 reconstructs the outer furry mesh, Gaussian Frosting supplies local fur-thickness cues: the method wraps a base mesh in an adaptive layer of 3D Gaussians, and where that shell is wide, more volumetric rendering was needed, which correlates with thick fur. Part labels come from ALIGN-Parts, a direct 3D part segmentation, so no SMAL fitting is required and non-quadruped mammals are in scope. Shell widths act as evidence, part-level priors calibrate them, and a bounded smoothing optimization moves each vertex inward along its normal, with non-fur regions locked and displacement capped near opposing surfaces to avoid flips and self-intersections. Strand length is initialized separately from defurring, since a curved strand is longer than the coat is thick: each part takes its 75th-percentile shell width times a part-specific multiplier, and an optional variant blends in 40 percent of a VLM length estimate.

Stage two is where the speedup comes from. Instead of optimizing tens of 3D control points per strand, an encoder maps each root's position, part label, local shell thickness, and target length to a compact vector of PCA coefficients. A decoder initialized from PERM, a PCA basis learned on human hair, turns those coefficients into a normalized strand in a local frame, scaled to target length and oriented with the root's tangent frame. Cylindrical Gaussians attached to the strand segments plug into 3DGS differentiable rendering, and the whole system trains end to end against photometric, silhouette, orientation (Gabor), anti-penetration, and other losses, fine-tuning the PCA decoder along the way. A human-hair prior transfers across species with no animal-fur training data at all, sampling 15K strands per iteration for 2,500 iterations.

Results

Same A5000 GPU, same 36 views:

MethodStrand trainingPreprocessing
NeuralFur10.5 h10 h
FurE52 min1 h

Roughly 10x faster strand training, and the full pipeline drops from about 20 hours to under 2. Geometry against an artist-authored ground-truth groom on a synthetic tiger (4 cm / 40 degree threshold):

MethodPrecisionF-score
GaussianHairCut32.3437.93
NeuralFur48.0546.84
FurE51.2050.86

Novel-view rendering on four Artemis scenes (panda, white tiger, fox, cat) is a three-way tie: panda PSNR 43.91 for both FurE and NeuralFur, cat 49.82 vs 49.84. The honest summary is faster, with equal rendering quality and slightly better geometry. Porting the human-hair method GaussianHairCut over fails visibly: mean strand length 8.79 on the panda versus 5.22 for FurE, and direction variance 0.56 versus 0.046, confirming that fur needs explicit per-part length modeling.

On a real captured multi-view bison sequence, NeuralFur cannot recover the underlying geometry at all while FurE produces editable strands. The authors call this the first instance-specific strand-based fur reconstruction from noisy real-world multi-view images, and they released the data. Ablations support both design choices: one global strand length misses the long fur on body and belly, and removing defurring degrades direction and curvature consistency.

Why it matters

For 3D capture and digital-asset work, per-instance fur reconstruction under two hours total is just reaching a practical threshold; twenty hours never was. The output is explicit strands that import directly into Blender and Unreal Engine, and the paper shows physically plausible motion under strong wind simulation. The recipe of borrowing a PCA basis from a data-rich domain instead of collecting data in the target domain travels well beyond fur: the basis captures local curve structure, not species identity. Dropping SMAL also extends coverage beyond quadrupeds.

Limitations

From the authors: the method needs calibrated multi-view images of a static animal, which living animals rarely cooperate with; defurring depends on Frosting cues and dense viewpoints, so heavy self-occlusion degrades both thickness and length estimates; fur is modeled as a single layer on a root surface fixed before optimization, leaving multi-layer coats and joint root-strand optimization open; there is no ground truth with strong within-part variation such as shaved patches or injuries; and metric scale comes from an assumed eye separation with no validated automatic fallback.

Two things deserve discounting. Rendering metrics differ in the second or third decimal, which shows parity, not superiority; speed is the real claim, and it holds. The length calibration is full of hand-tuned constants (75th-percentile shell width, per-part multipliers, a 60/40 blend, a 35 to 135 percent clip), the V2 variant leans on VLM estimates, and whether those survive a wider range of species is untested. The ground-truth evaluation covers a single synthetic tiger.

Terms

Source

Related papers

All paper explainers