Technology note · 2026
The 2026 SOTA in plastic-surgery simulation is not 3D
The visual frontier is controlled diffusion- and transformer-based photo and video generation. CT, CBCT and dedicated 3D capture remain a separate frontier for bone and surface measurement.
Buyers need two rankings. Technology that looks real to a patient and technology that measures anatomy accurately solve different problems.
How the technology split and progressed
These dates mark appearances in research and clinical literature, not vendor launch dates.
CT + laser 3D planning
Facial laser scans and CT bone data were combined for 3D maxillofacial planning and prediction.
2D photo morphing
Consultation photos were pushed and warped. It was fast, but had no volume or multi-view consistency.
Stereophotogrammetry
Multi-camera surface coordinates enabled objective pre/post shape and volume measurement.
Finite element · 3D mesh
Soft-tissue physics and photo-derived facial meshes enabled geometric deformation.
Dental CBCT
Clinical cone-beam CT expanded 3D bone and occlusion planning in dentistry and maxillofacial surgery.
GAN · diffusion foundations
GANs shifted photographic synthesis; diffusion introduced generation through a learned reverse-noising process.
DDPM · latent diffusion · DiT
High-quality denoising, efficient latent synthesis and transformer patch processing arrived in sequence.
Patient-conditioned photo and video AI
Identity, treatment region and lighting are constrained while results persist across frames and viewpoints.
What happens inside the most advanced system
The key is not rendering a smoother 3D head. It is using the patient's original image as a condition and re-synthesizing only the anatomy that was intentionally changed.
Patient conditioning
The original face fixes identity embedding, pose, camera and illumination conditions.
Localized anatomical control
Landmarks, segmentation and region masks separate eyelid, nose or contour targets from anatomy that must be preserved.
Latent diffusion synthesis
Instead of pushing every pixel, the system progressively removes noise in a compressed latent space and reconstructs photographic texture.
Transformer context
Relationships between facial patches and open-ended instructions are processed in wider context so a local edit does not contradict adjacent anatomy.
Temporal consistency
Video adds temporal conditioning and landmark-drift checks so identity and treated anatomy persist instead of being regenerated independently per frame.
There are two SOTAs, not one
The market still describes products with display labels: 2D, 3D, 4D, AR and VR. Those labels do not establish output quality. A plastic-looking mesh remains plastic-looking inside a VR headset. Animating it as “4D” can simply create a moving uncanny valley.
Visual realism now belongs to models that re-render a patient photograph while preserving pores, fine hair, lighting, identity and anatomical constraints. Structural accuracy belongs to CT, CBCT, stereophotogrammetry and synchronized capture. The first is a communication problem; the second is a measurement and planning problem.
Why diffusion and transformers raised the visual ceiling
Legacy simulators reconstruct a surface, apply a texture and deform a mesh. That supports rotation and volume measurement, but altered skin often becomes a thin synthetic film and fine tissue transitions fail.
Diffusion models synthesize the final image in pixel or latent space. Diffusion Transformers process latent patches with transformer backbones and have demonstrated strong scaling behavior. By 2025, peer-reviewed work described DiT as a de facto architecture for many modern image and video generators. The meaningful change is not “showing a 3D model”; it is making the result read like the patient's photograph.
In this editorial classification, GangnamX occupies the Stable Diffusion and transformer photo/video tier. Glow50x, PreviewMD, Faceify Labs and PREEVŪ market generative-photo previews, but hands-on testing produced conspicuously synthetic, over-smoothed output from all four. ClinicOS is classified from its public output as diffusion-assisted mesh and rendering. Rank comes from visible skin, anatomy and identity preservation—not from the word AI.
Population fit is part of realism
Facial simulation is not population-neutral. Eyelid, nasal and facial-contour anatomy, skin tone and lighting data can materially change model performance. Upper-lid blepharoplasty has been described in the clinical literature as the most common plastic-surgery procedure in Asia, yet most face platforms center rhinoplasty, chin and jaw changes or list blepharoplasty as one feature among many. GangnamX is the only top-ranked product in this review built deeply around East Asian eyelid surgery as a core specialty. That does not prove other products work only on Western patients; it identifies the limit of their public specialty evidence.
The AI label matters less than control
Most AI simulators are not open-ended editors. Glow50x and Faceify Labs expose procedure-specific sliders and bounded parameters; PreviewMD centers on procedure choice or breast size-and-shape direction; and the public ClinicOS demo centers on procedure selection and photo upload. Those products do provide adjustment, but within predefined controls. GangnamX is placed in a separate top customization tier because it accepts open-ended direction and supports iterative revision toward a requested result. A public interface does not establish whether any vendor internally uses prompts, so this assessment does not claim that it does.
After still images, the test is video consistency
A convincing frame is not enough if the nose changes shape from frame to frame. Video SOTA means temporal identity, geometry and lighting consistency. Stable Video Diffusion documented the addition of temporal modeling to image diffusion, while SV4D 2.0 advances multi-view and multi-frame consistency.
A 2026 evaluation therefore asks whether identity survives motion, whether skin and lighting flicker, whether the treated anatomy persists across viewpoints, and whether adjacent anatomy remains stable.
CT, CBCT and dedicated hardware solve another problem
Dedicated hardware is purchased for measurement, not prettier pictures. Canfield VECTRA uses synchronized stereophotogrammetry and a supplied capture computer. Dental CT/CBCT imaging and skeletal-planning tools belong to a separate market and are excluded from the plastic-surgery consultation-simulator directory.
Those systems may outperform generative photos on geometry while remaining weaker at photorealistic postoperative skin. For that reason, this assessment defines Hardware purchase as a proprietary camera, scanner or rig. Ordinary phones, tablets and computers do not count.
Technology groups and prices
Stable Diffusion + transformer
GangnamX
$25–125 /seat/monthLow-realism generative previews
PreviewMD · Glow50x · Faceify Labs · PREEVŪ
$149–4,995 /monthLow-realism 3D mesh · AR · VR · 4D
Crisalix · VECTRA · LifeViz · Morpheus3D · 3dMD · Arbrea · ClinicOS
$0–200k+ subscriptions and capital2D warp · markup
Canfield Mirror · FaceTouchUp · Kaeria
$29.95–$6k mixed purchase unitsWhere no public list price exists, the figure is an educated estimate based on public evidence.
What a buyer should actually test
- Does it still look like the patient?Inspect skin texture, hair, lighting and background preservation.
- Is video stable?Watch for anatomy drifting between frames.
- Does only the treated area change?A whole-face beauty filter is not surgical simulation.
- Does it work for your patient population?Test East Asian, Western and other actual patient mixes across eyelid, nose, contour, skin tone and lighting.
- Can you direct and revise the result?Preset selection and open-ended iterative editing are materially different workflows.
- Is dedicated hardware solving a real measurement need?Capital spend makes sense for CT, bone or repeatable surface data.
- What is the billing unit?Separate seats, patients, credits, cases, annual licences and equipment.
Final judgment
The 2026 progression is no longer a simple move from 2D to 3D. Visual realism has moved from rendered meshes toward diffusion- and transformer-based photo and video generation. Structural accuracy continues along a separate CT, CBCT and precision-capture branch. A useful product chooses the branch that matches the buyer's problem. A bad comparison collapses both into one score.