NEO 2: 40% Faster Generation with Latent Cascading
NEO 2 now generates avatar video about 40% faster, with essentially no drop in visual quality. It does the expensive work on a grid half the width and half the height, then restores full resolution in a separate finishing pass.
Swipe across the diagram, or open it at full size below.
The Latent Cascade
The expensive part of generating a NEO 2 video is the main diffusion transformer, run over and over. Every denoising step asks it to process a spatiotemporal volume, and increasing the spatial resolution expands that workload at every one of them. So we changed where that compute goes: the main DiT generates at half the latent height and width, and a dedicated finishing path restores the target latent resolution before decoding.
We call this two-scale path the Latent Cascade. It combines Coarse Generation, a reference-conditioned upscaler we call Reference Lift, and Flow Refinement. Together, they separate the expensive task of generating an avatar performance from the task of reconstructing it at the output scale.
In our internal testing, the Latent Cascade delivered an overall speedup of approximately 40%. 49% of participants preferred the upscaler version, 34% preferred the previous version, and 17% rated them a tie. Faster, with essentially no loss in perceived visual quality.
01 / Two spatial scales
Keep the temporal canvas. Reduce the spatial workload.
The saving comes from one decision: generate on a smaller grid. NEO 2 already generates in the compressed representation of a video VAE, and the cascade introduces a second spatial scale inside that latent space.
In the target-scale path, the main DiT denoises latents at the spatial resolution required by the final decoder. In the cascade, it instead operates on a volume with half as many positions along each spatial axis. The latent temporal length stays the same. With a fixed patch geometry, halving both spatial dimensions reduces the spatial token count to one quarter, apart from padding effects.
The reference image follows two conditioning paths. A half-size image is VAE-encoded for Coarse Generation, while a target-size encoding supplies the finishing stages. This gives the lower-resolution generator a compatible reference and gives Reference Lift and Flow Refinement access to target-scale appearance information.
The diagrams depict each latent video as a three-dimensional lattice over time, height, and width. The actual tensor is five-dimensional: batch × channels × time × height × width. Every lattice site contains a 16-channel feature vector; depth in the drawings represents time, rather than channels.
Swipe across the diagram, or open it at full size below.
In plain terms: the video keeps every frame it started with, and only the picture inside each frame grows. H and W denote the target latent dimensions, rather than output pixels. The finishing path preserves T; the VAE subsequently reconstructs the pixel video.
02 / Inside the cascade
Three stages, and only one of them repeats.
Generating on a smaller grid only works if the finishing stages can put back what that grid left out. The main generator establishes the performance. Reference Lift expands the latent grid with spatiotemporal processing, and Flow Refinement applies a timestep-conditioned correction.
Coarse Generation
This is the stage the speedup comes from. The existing NEO 2 DiT receives audio, text, and the lower-resolution reference. It performs repeated denoising on the compact grid to produce z↓. The reviewed L40S inference preset uses three main-model steps; this schedule is a configuration choice, rather than a requirement of the cascade.
Reference Lift
This is the stage that puts the resolution back. A learned 1×1×1 projection and spatial channel rearrangement double the latent height and width. The expanded volume is concatenated with the target-scale reference, broadcast across time. A stack of 3D residual convolutional blocks and temporal attention predicts a correction that is added to the expanded latent.
Flow Refinement
A short correction pass over the lifted result. A separate DiT takes the lifted latent, the target-scale reference, a reference-frame mask, and a timestep. It predicts a velocity field that the flow-matching scheduler uses to update the latent. The reviewed preset mixes in a small amount of noise before one refinement step.
Overlapped output
Decoding and encoding run at the same time instead of one after the other. The VAE decodes the finished latents into successive RGB batches. Host transfers and a bounded producer–consumer queue feed a video-encoding thread, allowing encoding to proceed while decoding continues. The worker returns the encoded bytes once encoding is complete.
Swipe across the diagram, or open it at full size below.
Technical notes: residual lifting, temporal attention, and refinement inputs
Reference Lift first computes u = S₂(z↓), where S₂ is the learned spatial sub-pixel expansion. It then computes z↑ = u + F([u, r↑]), where brackets denote channel concatenation and r↑ is broadcast across T. This explicit skip connection makes the learned branch a residual correction to the expanded latent. Temporal attention operates along T at each spatial site, with rotary positional embeddings on the temporal queries and keys.
The upscaler class defaults to 16 residual blocks, with temporal attention inserted after every fourth block. The refiner class defaults to (1, 2, 2) patch embedding, three-axis rotary positions, timestep modulation, and alternating shifted spatial attention windows that span the input's temporal extent. These are source-code defaults; the loaded checkpoint determines the instantiated architecture.
The refiner concatenates 16 video channels, 16 reference channels, and one mask channel. The mask is zero for the leading reference frame and one thereafter. Its forward pass receives no fresh audio or text conditioning; that information has already influenced the generated video latents.
In the reviewed L40S preset, the upscaler noise coefficient is zero and the refiner noise coefficient is 0.1, giving z̃ = 0.9z↑ + 0.1ε. The worker then applies the configured flow update, z* = z̃ + (σnext − σ)vθ. The noise coefficient and scheduler timestep are separate settings.
Keep reconstruction moving.
The output path uses a different form of efficiency: overlap between decoding, transfers, and encoding. On CUDA, asynchronous copies use pinned host buffers and synchronization events. The bounded queue applies backpressure if encoding falls behind. The decoder also removes the leading reference frame before yielding the generated RGB frames.
This is internal streaming between pipeline stages. It does not, by itself, establish a progressively playable response or a shorter user-visible time to first frame.
Swipe across the diagram, or open it at full size below.
03 / Internal evaluation
Faster inference, with visual quality maintained.
In our internal testing, latent cascading delivered an overall speedup of approximately 40% over the previous NEO 2 inference path. The visual-preference results favored the upscaler version.
- Overall inference speedup
- ~40%
Compared with the previous NEO 2 inference path in internal testing.
Which version did participants prefer?
- Preferred the upscaler version
- 49%
- Preferred the previous version
- 34%
- Rated them a tie
- 17%
49% of participants preferred the upscaler version, while 34% preferred the previous version and 17% rated the two versions a tie. These results support the central tradeoff of the Latent Cascade: faster inference with essentially no loss in perceived visual quality in our internal evaluation.
See the two versions side by side.
Compare the previous NEO 2 version on the left with the Latent Cascade output on the right. The clips are synchronized in one video; use full-screen playback to inspect the detail.
04 / The compute tradeoff
Spend repeated compute on the compact grid.
We moved the main DiT's repeated denoising onto a smaller spatial lattice, and paid for it with one learned lift and a short refinement schedule at the target scale. This is a multiscale allocation of inference compute, building on the latent generation described in the original NEO 2 article.
A fourfold reduction in spatial token count does not imply a fourfold end-to-end speedup. Different model operations scale differently with sequence length, and the finishing stages introduce their own cost. Input preparation and output reconstruction also contribute to the request's critical path. Because decoding and encoding overlap, their separate durations should not simply be added as if they ran sequentially.
The target-scale reference provides appearance information for reconstruction, while temporal operators let the learned correction depend on neighboring latent frames. Our internal preference results support this design: more participants preferred the upscaler version than the previous version. The approximately 40% overall speedup therefore comes with essentially no perceived quality loss in that evaluation.
By shifting repeated generation to a compact latent grid, the Latent Cascade delivers an approximately 40% overall inference speedup. Internal testing indicates that this gain comes with essentially no loss in perceived visual quality. The expensive work now happens on a quarter of the spatial sites. The rest is put back.
Frequently Asked Questions
The latents occupy a three-dimensional spatiotemporal lattice: T × H × W. The full tensor also includes batch and feature-channel dimensions, making it B × 16 × T × H × W. The diagrams show the lattice geometry and leave those two additional dimensions implicit.
The latent upscaler expands only the spatial dimensions. It preserves the input latent temporal length, so this change does not downsample the generated sequence in time. Final frame reconstruction remains the VAE's job.
No. Reference Lift begins with learned spatial sub-pixel expansion, then applies a reference-conditioned residual network with 3D convolutions and temporal attention. Flow Refinement is a separate stage that predicts a velocity update on the expanded latent volume.
No. The main NEO 2 DiT still runs its own denoising schedule. The reviewed L40S preset uses three main-model steps, followed by Reference Lift and one refiner step. These counts describe that preset and can change with configuration.
Our internal testing showed an overall speedup of approximately 40% compared with the previous NEO 2 inference path.
In our internal testing, 49% of participants preferred the upscaler version, 34% preferred the previous version, and 17% rated them a tie. These results suggest that the speed gain comes with essentially no loss in perceived visual quality.
Related Research
NEO 1.5: Identity-Aware Performance
Speaker identity moves from the side of the network into the middle. Inside a single Unified Joint Attention pass, audio, motion, and identity reason together — fixing the moments lip-sync usually breaks.
Beyond Lip-Sync: Robust Video Dubbing
Most lip-sync models break the moment conditions get difficult. Colossyan Dubbing handles real-world footage by feeding the full frame through NEO 2 instead of cropping the mouth, producing broadcast-quality output on raw, unprocessed video.
NEO 2: Full-Body Expressive AI Avatars
The model that powers Colossyan Dubbing under the hood. Full-body, audio-driven avatar generation with consistent identity at any duration.
NEO: Expressive Talking Head Performance
The first generation of NEO. Talking-head performances with natural head movements, facial expressions, and eye gaze from audio input.


