Vestra
Scene contract, reconstruction, fusion, exports, and browser studio. The product owns the world and its provenance.
scene / studio02 / SPATIAL COMPUTING CASE STUDY
Vestra turns handheld RGB video into an inspectable relative-scale scene — with the hard parts left visible: provenance, registration, residuals, and the limits of what the data can support. The public demo is TUM RGB-D freiburg1_room (CC BY 4.0), reconstructed without ground-truth poses.
01 / THE PROBLEM
A per-frame depth map can look convincing and still fail to describe one coherent room. Cameras drift. Scale changes between windows. A dense preview can hide broken registration. Vestra treats reconstruction as a chain of evidence: every derived surface stays attributable to the input and the transform that produced it.
02 / THREE REPOSITORIES
One product boundary, three explicit ownership boundaries.
Scene contract, reconstruction, fusion, exports, and browser studio. The product owns the world and its provenance.
scene / studioNative Rust inference runtime, ported from and benchmarked against depth-anything.cpp (localai-org, C++17/ggml): calibrated preprocessing, depth, confidence, and camera pose. The Depth Anything 3 model itself is upstream work by ByteDance.
inference / CPUShape-gated low-level CPU and experimental CUDA kernels, qualified independently before adoption.
kernels / parity03 / IMMUTABLE EVIDENCE
Raw measurements, derived geometry, and generated presentation are intentionally separate layers. A beautiful render cannot silently rewrite what was measured.
04 / MEASURED, NOT MARKETED
Every public measurement must identify the exact source revisions, workload, and configuration it describes. The older figures have been removed: they are not being presented as measurements of the pinned portfolio snapshot.
Quantitative performance is under revalidation against the exact pinned Vestra Engine and Vestra Kernels revisions. Until that check is complete, this page makes no numeric performance claim for the public product release. The pinned sources are linked above.
05 / INPUT AND INSPECTION
Two silent recordings from the same public TUM room fixture: a short excerpt of the RGB input and Vestra's browser viewer orbiting a frozen, precomputed scene. They are not synchronized, and playback speed says nothing about reconstruction time.
A nine-second excerpt from the canonical 45.4-second TUM RGB-D freiburg1_room RGB input. The excerpt is re-encoded for web playback without cropping or geometric edits. Source: TUM RGB-D Benchmark, CC BY 4.0.
Vestra renders a 355,581-point geometric-MVS control derived by the pinned COLMAP provider and imported under the scene contract. The scene links 55 selected source frames; it uses no ground-truth camera poses and is not presented as pure Rust. The local reconstruction of the same 20 Aug 2026 demo record (4,275,936 measured points / 9 windows) is a separate ledger line, not this clip.
06 / HONEST LIMITS
01Relative scale is the current truth. A handheld video does not become metric just because its point cloud is dense.
02There is no semantic mesh claim. Surfels and measured geometry are useful without inventing object labels.
03COLMAP is an optional, pinned provider for global registration experiments — not a pure-Rust implementation.
04A rejected pose hypothesis remains rejected evidence. Rendering more points cannot repair a broken global trajectory.
05Inference, reconstruction and the viewer are separate stages. The orbit clip is display frame rate on a frozen scene, not a reconstruction timing.
06No public binary demo download; the v0.1.0 binary bundle was withdrawn on 30 Aug 2026. Source is public, the app runs locally, and there is no hosted instance or uptime claim.
07Scene publishing is atomic against process abort (write, then rename). It is not fsync-durable against power loss.
08No speed number on this page is treated as a property of the product until it is verified against the pinned engine and kernel revisions.
THE TAKEAWAY