Super-resolution ×4 — SeedVR2 vs Real-ESRGAN (Set14, RTX 5090)

Timings & VRAM

VariantCFGStepsLoad Avg gen / imgPeak VRAMPSNR ↑SSIM ↑LPIPS ↓DISTS ↓NIQE ↓MUSIQ ↑CLIP-IQA ↑Success
Bicubic x4 (baseline) 0.0s 24.060.66880.44250.22407.6332.420.4938 14 / 14
Ground truth (original HR) 4.94°66.94°0.7418° 14 / 14
Real-ESRGAN x4 (fl16) 1 3.9s 0.1s 0.25 GB22.480.62200.24600.16474.7565.090.6247 14 / 14
SeedVR2 3B (bf16, 1-step) 1.0 1 7.6s 0.7s 8.11 GB21.390.56640.22930.14753.9071.990.7565 14 / 14
SeedVR2 7B (bf16, 1-step) 1.0 1 9.2s 0.9s 17.70 GB22.060.60300.18770.13423.9772.250.7804 14 / 14
SeedVR2 7B-sharp (bf16, 1-step) 1.0 1 9.2s 0.8s 17.70 GB21.810.59270.19190.13003.8372.340.8074 14 / 14

Shaded columns are no-reference — they judge “does this look like a real photograph?” with no ground truth. They have no winner: the ° ground-truth row is the calibration line, and out-scoring it means inventing texture prettier than the original. The unshaded columns are full-reference and measure distance to the ground truth. A model can win one axis and lose the other — that is the perception–distortion tradeoff, not a contradiction.

Drag the slider to compare A vs B; switch either dropdown to pick any variant. The number next to each dropdown is that variant's generation time for this example. Click the input thumbnail to view it full size.

example #baboon

input:
vs

example #barbara

input:
vs

example #bridge

input:
vs

example #coastguard

input:
vs

example #comic

input:
vs

example #face

input:
vs

example #flowers

input:
vs

example #foreman

input:
vs

example #lenna

input:
vs

example #man

input:
vs

example #monarch

input:
vs

example #pepper

input:
vs

example #ppt3

input:
vs

example #zebra

input:
vs