An ongoing attempt to reconstruct the 1970s Aspen Movie Map as a hierarchical 3D Gaussian Splat scene, built from a 16 mm film scan of the original MIT footage.
Work for a Member organization and need a Member Portal account? Register here with your official email address.
An ongoing attempt to reconstruct the 1970s Aspen Movie Map as a hierarchical 3D Gaussian Splat scene, built from a 16 mm film scan of the original MIT footage.
Aspen Movie Map was a project at the MIT Architecture Machine Group, led by Andy Lippman, in the late 1970s. A gyro-stabilized car drove every street in Aspen, Colorado with four 16 mm cameras pointing north, south, east, and west, snapping a frame every ten feet. The frames were pressed onto Laserdisc and indexed so that a viewer could pick an intersection, pick a direction, and have the disc cue up the next clip. It is one of the earliest hypermedia systems and a direct ancestor of Google Street View, about twenty-five years before Street View existed. MIT Media Lab posted a short overview on YouTube.
A Laserdisc copy of the project lives on the Internet Archive as ASPEN4, the version of the software Andy used in his later demos. You can scrub through it and watch an Aspen winter from inside a 1978 station wagon.
There's a personal thread here. As an undergrad I worked for Michael Naimark, who ran cinematography on the original Aspen Movie Map. He is a media artist whose work has been about place representation for forty years, and he is the reason I knew this project well enough to want to bring it back.
Wouldn't it be nice to reconstruct the Aspen of the 1970s in the full 3D glory of the 2020s?
We have Gaussian Splatting now. But not just any Gaussian Splatting. The dataset is big. Four cameras, every ten feet, every street in town, hundreds of thousands of usable frames. Vanilla 3DGS cannot hold a scene like that in memory.
For that I am using A Hierarchical 3D Gaussian Representation for Real-Time Rendering of Very Large Datasets from Inria. It is built for exactly this regime: city-scale captures, level-of-detail rendering, splats organized into a hierarchy that streams as you move.
I started with what was easiest to find. A Laserdisc rip of the project. I pulled the driving sequences, ran COLMAP on them, and brought the result into Agisoft Metashape. It did not work. Laserdisc is composite NTSC, lossy, interlaced, and noisy. Edges shimmer, color bleeds, feature matching collapses on most frames.
I emailed Andy. He told me that Cinelab in Massachusetts had scanned the original 16 mm reels at high resolution within the past few years. I walked to his office, picked up a drive, and copied a few hundred gigabytes of ProRes transfers off it.
I reran COLMAP on the driving sections of the film scan. Out of about 36,000 frames, 27,000 aligned in Agisoft Metashape. The sparse point cloud already carries the shape of the Aspen street grid and the contour of the mountain rising on the south side of town.
I tried vanilla 3D Gaussian Splatting first, the original Kerbl et al. (SIGGRAPH 2023) implementation. It blew up my memory, as it should. Vanilla 3DGS holds the entire scene in VRAM, and an entire town does not fit.
So I switched to the hierarchical pipeline. A Hierarchical 3D Gaussian Representation for Real-Time Rendering of Very Large Datasets is by the same Inria GraphDeco / FUNGRAPH team that wrote the original 3DGS paper. It is built for exactly this regime: divide the scene into chunks, train each chunk on its own subset of cameras, then merge the trained chunks into a single hierarchy at render time.
The pipeline divided my scene into 48 chunks. Each chunk gets trained individually for 30,000 iterations, the default in the paper. On my laptop 4090, that takes days and days. I let it run for about a week.
A side effect of the chunking step is a per-frame depth map, generated by Depth Anything v2 inside the hierarchical pipeline. Not strictly important for the project, but it has an interesting look, knowing this is the street from the 70s.
I previewed individual chunks in Postshot before the full hierarchy was done. From the original camera path, where the car drove, the reconstruction holds. You can recognize buildings, signs, sidewalks, and the angle of the road. Off the path it falls apart, the same way every Gaussian splat does at the edge of its sample.
The two artifacts I see most are predictable:
After a week of training I tried to load all 48 chunks together in the hierarchical viewer. My 4090 laptop cannot do one frame per second. The PLY files alone, before the viewer even loads them, are larger than my 16 GB of VRAM.
We are setting up an RTX A6000 on an external GPU enclosure for the laptop. 48 GB of VRAM should clear the loading problem, and the A6000 should be fast enough to render the merged hierarchy at a reasonable frame rate.
This post is a work in progress. The training is done, the chunks look right individually, but the merged Aspen does not run on the hardware in front of me yet. Next update is when the A6000 works and the full city loads.