The Symplectic Camel
Gromov's non-squeezing theorem, shown with four live WebGL scenes. A volume-preserving map can squeeze a 4-dimensional ball through a hole narrower than the ball. A symplectic map, the kind classical mechanics allows, cannot, and the film measures the ball's shadow every frame to show it.
Making of
The mathematical idea
A ball in R^4 can be squeezed through a hole narrower than itself if we only demand that volume be preserved: shrink the first plane, stretch the second. Gromov's 1985 non-squeezing theorem says that a symplectic map cannot do this. If the ball B^4(r) embeds symplectically in the cylinder B^2(R) x R^2, then r <= R. The point is that symplectic maps, which are the maps of Hamiltonian mechanics, preserve more than volume: they remember area in the (x, y) planes.
The film uses the four coordinates as phase space for a system with two degrees of freedom, so the idea reaches mechanics. Preserving the symplectic form forces y1 to stretch when x1 is squeezed; a nonlinear Hamiltonian flow may twist and fold the ball, but the shadow on the (x1, y1) plane never drops below pi r^2. The last scene turns the cylinder so that its base is the Lagrangian (x1, x2) plane, two positions and no momentum, and then the ball fits.
How the visual was built
The scenes are WebGL (Three.js) animations of a 4-ball sampled as a point cloud, with the fourth coordinate shown by colour. The shadow on the (x1, y1) plane is measured from the points every frame and plotted live, so the viewer can watch the number stay above the floor.
The film is made by a deterministic pipeline in the project's video directory. A script file holds each line of narration with its on-screen text and a spoken respelling. A local text-to-speech model (VoxCPM2) generates one clip per line, keyed by a hash of the text, and a local Whisper model re-reads every clip to check the words. An assembly step builds the timing file and the subtitles. The picture is a page with a function renderAt(t), so any frame can be reproduced exactly. A capture script drives system Chrome through puppeteer-core, pipes PNG frames to x264, and muxes the audio.
One encoding lesson is worth recording. The file size was determined by the source, not the ffmpeg flags: sparse one-pixel point clouds plateau around VMAF 91 even at 37 Mbps. Rendering at 2x supersampling with 320,000 fainter points reached VMAF 94 at 6 Mbps, and avoiding a JPEG intermediate mattered.
What is verified and what is not
Nothing here is formally verified. The theorem is cited, not proved. What is checked is the measurement: the shadow area is computed from the actual points in each frame. The 4D-to-3D projection and the particular flows are choices made for the picture, and the labels on screen say which scene is which.
By the numbers
- Selected evidence: 3 sessions (2 Claude main, 1 Codex desktop), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 12 direct human-message prompts, 0 queued typed prompts, and 3 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 103 logged agent output units; 93 tool calls; 1 new Agent requests. Output units differ between formats. Claude logged 130,413 output tokens.
- Models recorded: claude-opus-5-5, gpt-6-sol. Core records cover 1 active days, spanning 2.2 elapsed hours, including waits. The current spoken script has 55 lines and 560 words, counted from its text fields.
In Ben's words
“I particularly liked Lagrangian cylinder and Hamiltonian flow.”
2026-09-25
“pi is pronounced incorrectly twice and correctly once so far”
2026-09-25
“capital R whenever it's.. R would be better”
2026-09-25
How the AI collaboration went
The main thread moves from an open choice of visualization tools to a narrated film in the same evening. Ben first asked for a performant animation, then selected the Hamiltonian-flow and Lagrangian-cylinder scenes by reacting to the browser result. His next turn changed the deliverable to a video with voiceover and restrained subtitles. Claude supplied the implementation and the production scripts; Ben supplied the selection and the listening corrections.
The first export drew an immediate objection: Ben reported a 945 MB file. A retained command compares seven encoded test clips against a reference using VMAF. That is seven comparison candidates, not seven complete film renders. The useful fix was to change the source image: denser, fainter points and supersampling made the cloud easier to compress. Later, Ben heard “pi” read inconsistently and noticed that the two radius letters were not distinguished. Claude generated replacement voice takes and reused the picture. These repairs explain why a deterministic, separately timed picture and audio track mattered. A visually finished export still needed mathematical pronunciation and presentation judgment.
Credits
Made by Ben Knill with Claude Code (Claude Opus). Narration: synthetic voice (VoxCPM2). Three.js from cdnjs. Page and source: github.com/BenKnill/symplectic-camel.
Chapters
- 0:00 A ball and a narrower cylinder
- 0:30 Four dimensions are phase space
- 1:06 How to read the picture
- 1:40 A map that is not symplectic
- 2:04 A linear symplectic squeeze
- 2:28 A nonlinear Hamiltonian flow
- 3:04 Gromov's theorem
- 3:36 A Lagrangian cylinder
- 4:07 What a symplectic map remembers
- 4:42 Outro
Running Chaos Backwards
A kicked rotor (Chirikov's standard map) scrambles a picture of a camel into noise. Run backwards in floating point, only the islands of stability come home. On a 256 x 256 lattice each step is two whole-cell shears, so it permutes the cells and every pixel returns. The 92 bytes of machine code that do it are proved correct in HOL Light.
Making of
The mathematical idea
The kicked rotor, Chirikov's standard map, is a model of chaos: nearby orbits separate exponentially. It is also symplectic, so in principle it can be run backwards exactly. In floating point it cannot. Each step amplifies the rounding error, so a few dozen steps in, the computed inverse no longer returns the picture; only the islands of stability, where orbits are not chaotic, come home.
The fix is to move the world onto a 256 x 256 lattice. Writing the map as two shears, and rounding the kick to a whole number of cells, makes each step a permutation of cells. A permutation has an exact inverse, so every pixel returns, however chaotic the forward map is.
How the visual was built
The film uses the same deterministic pipeline as episode 1: a script file, a local text-to-speech model, ASR re-reading of every clip, a film page with a function renderAt(t), and headless Chrome capture to x264. The picture that is scrambled is the camel from episode 1. The file is 4:08 at 1080p and only about 70 MB, because the blocky three-by-three cells compress very well.
A companion lab page (Do, Undo, Repeat) puts four cards on one shared ten-cycle clock: chaos, a 30 degree rotation done by bilinear resampling versus three exact shears, the Vancouver Stock Exchange index of 1982, and the Patriot missile clock of 1991. The Patriot card reproduces the published GAO figure by chopping 1/10 to 23 fractional bits.
What is verified and what is not
The forward and backward steps exist as real AArch64 machine code, 92 bytes in all. Their correctness and the round-trip theorem are proved in HOL Light for any kick table and any n below 2^63, using the s2n-bignum ARM model, and the Hearth replay completed with no new axioms. A NEON version of the forward step, about 2.9 times faster, is proved against the same statement, which was checked term for term.
Everything outside that is checked or illustrative. The JavaScript port that draws the page matches the verified kernel at all 201 film states. The floating-point lane is ordinary double arithmetic, shown as the foil. The loop that calls the kernel, the float comparison and the drawing are not covered by the proof.
By the numbers
- Selected evidence: 3 sessions (3 Claude main), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 5 direct human-message prompts, 0 queued typed prompts, and 7 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 308 logged agent output units; 309 tool calls; 0 new Agent requests. Output units differ between formats. Claude logged 632,062 output tokens.
- Models recorded: claude-fable-5-1, claude-opus-4-8, claude-opus-5-5. Core records cover 3 active days, spanning 35.6 elapsed hours, including waits. The current spoken script has 42 lines and 486 words, counted from its text fields.
In Ben's words
“I like the scramble-unscramble demo. Take it to the next level.”
2026-09-26
“we aren't prepublication so there's no from-scratch check needed”
2026-09-26
“don't need cold replay except prior to publication (meaning external publication).”
2026-09-27
How the AI collaboration went
This grew out of Ben's question about using verified assembly to visualize dynamics. The first useful object was a reversible picture, rather than a general simulation platform. Claude built the lattice map, its inverse, the camel-image display and the film using the narration machinery from the preceding episode. Ben then asked for repeated scramble/unscramble cycles and less obvious examples of accumulated numerical damage. The four-card lab was a response to that request; it is a companion experiment, not part of the assembly theorem.
The optimization continued in a second Claude main session. Its opening is a proof handoff: lift the one-lane lookup lemma to sixteen lanes, establish the loop invariant and finish the scalar tail. That copied command is evidence of continuity, not a publishable quotation of Ben's original prose. Ben's own follow-up urged the stalled lane forward and distinguished warm development from a cold replay before external publication. The failure-and-repair trail is unusually specific: unsupported instruction forms, an array invariant broken at stores, expensive bit-blasting and over-broad rewrites. The eventual result was an optimization under the same contract. The film's error-free reversal and the faster kernel were connected by explicit comparisons rather than assumed to be the same computation.
Credits
Made by Ben Knill with Claude Code (Claude Opus). Narration: synthetic voice. Proofs: HOL Light with the s2n-bignum ARM model, checked through Hearth. The Vancouver and Patriot figures follow the cited public records (GAO). Page and source: github.com/BenKnill/lattice-echo.
Chapters
- 0:00 A picture about to be scrambled
- 0:38 The kicked rotor
- 1:26 Floating point forgets
- 2:02 A grid of 256 by 256 cells
- 2:49 The machine code and its proof
- 3:41 Last time: symplectic maps
- 4:01 Outro
Float Unscramble: where trust ends
A chaotic map scrambles a picture and is run backwards in 64-bit floats. A proved envelope marks the step count N up to which the round trip is guaranteed; the plain float build loses its first pixel at N = 31, just after the guaranteed horizon of N = 25.
Making of
The mathematical idea
Episode 2 ran a chaotic map backwards exactly by moving to integers. This film asks the opposite question: how far can you trust the floating-point version? The map is a bent cat map, f(x) = x + K G(x) with K = 1/4 and the cubic G(x) = x(1 - x)(1 - 2x), on the torus. It is hyperbolic everywhere. A worst-case error analysis gives an envelope that grows geometrically with the number of steps N; the film draws that envelope as a halo.
The envelope guarantees a round trip for N up to 25 on a 256 x 256 picture. The actual builds do better than the worst case, but not by much: the plain build and the fused-multiply-add build both lose their first pixel at N = 31, six steps past the horizon. A double-double reference, which is not binary64, holds out to N = 70.
How the visual was built
The film is pixel art, not a ray-traced render. A Python program with numpy and Pillow draws every frame from recorded states and pipes them to ffmpeg, which takes about two minutes on the bluestar machine. The middle and right panels are the proved object file itself (labelled plain build, and plain build plus proof); the left panel is a second, unproven build using fused multiply-add. Every state of all eleven round trips was hashed on an Apple M5 and again under qemu, and all 535 hashes agree, so the film uses the qemu states.
The halo is the theorem's own envelope function, evaluated at K = 1/4 and drawn in pixels. When it is below a pixel, a 7 x 7 window magnified 30 times shows it at true scale; beyond that it is a blur of the same radius with a blue tint over the right panel.
What is verified and what is not
Proved in HOL Light through Hearth, with the s2n-arm-fp64 profile and no new axioms: the forward and backward step kernels are correct as machine code (12 of 12 bindings, 125 of 125 specification pins matched), and the run loop that calls them is too (18 of 18 bindings, 136 of 136 pins). The horizon is evaluated inside HOL, one matrix application at a time, so "guaranteed for N up to 25" is a theorem and not the output of a script.
Two design choices made the proof possible. Every value stays in [0, 1] by the nearest-rounding property of round-to-nearest-even, not by exactness arguments. And the mod-1 reductions compile to compare-and-branch, so the proof splits on the comparison, giving twelve paths per loop body. The cat-map kick, rather than the cubic standard map, was chosen because the standard map has islands of stability and the worst-case envelope then promises only a quarter of what floats actually deliver.
Not proved: the drawing code, and the pixel-scale presentation of the halo.
By the numbers
- Selected evidence: 3 sessions (2 Claude main, 1 Claude subagent), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 0 direct human-message prompts, 0 queued typed prompts, and 5 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 252 logged agent output units; 264 tool calls; 1 new Agent requests. Output units differ between formats. Claude logged 13,422 output tokens.
- Models recorded: claude-opus-5-5. Core records cover 1 active days, spanning 3.3 elapsed hours, including waits. The current spoken script has 13 lines and 152 words, counted from its text fields.
- Retained proof receipts: 64 files (36 failed, 25 passed, 2 incomplete, 1 refused); includes probes and preparation, not distinct theorem attempts.
In Ben's words
“I'm interested in your thoughts on verified assembly simulators for visualizing interesting dynamics.”
2026-09-25 · shared series/workflow direction
“we're in the 'gather/make cool stuff to present' phase”
2026-10-05 · shared series/workflow direction
“If I could nudge in any direction it would be proven vs unproven side by side over long steps like scramble/unscramble”
2026-10-06 · shared series/workflow direction
How the AI collaboration went
The initiating creative instruction is shared with the sdot film: put proved and unproved computations side by side over a long run. The coordinator launched the two dedicated Opus agents at almost the same timestamp. For this film, the selected agent transcript spans about three hours and thirteen minutes on 6 October; that is elapsed transcript time, including waits, not a measure of continuous computation or Ben's working hours.
The agent had three concrete stages: a native experiment, a proof and a film. Its first mathematical choice did not tell the desired story. With the standard map, a worst-case guarantee of 27 steps sat far below the first observed failure at 123, because the stable islands complicated the comparison. It changed to an everywhere-hyperbolic bent cat map. The final guaranteed horizon of 25 and first observed pixel loss at 31 made the distinction visible without pretending that a worst-case theorem predicts typical failure exactly.
The proof then forced engineering choices: compare-and-branch reductions produced twelve paths per body, and range preservation used nearest rounding rather than an unsupported exactness argument. The retained status reports 184.4 seconds of step-proof evaluation and 66.6 seconds for the compiled loop. The renderer used recorded, cross-machine-hashed states and the theorem's envelope function. A two-dimensional compositor was enough for this pixel-scale question. The reliable direct prompt sample is sparse; the other quoted directions below concern the series and its workflow.
Credits
Made by Ben Knill with Claude Code (Claude Opus). Proofs in HOL Light via Hearth. Narration: synthetic voice.
Dimples on the Rhine
River dimples are the tops of whirlpools. Point vortices obey Kirchhoff's Hamiltonian equations, in which each vortex's x and y are a conjugate pair, so the river surface is its own phase space. Each dip lenses sunlight into a dark shadow with a bright caustic rim, and the film mixes our own footage from the High Rhine with a simulation.
Making of
The mathematical idea
In a flat, incompressible flow the vorticity is carried by the fluid, and a small concentrated patch of it behaves like a point vortex. Helmholtz's rule says vortex lines are carried with the fluid and cannot end in it, which is why a dimple at the surface must be the top of a tube that goes somewhere. For N point vortices the equations of motion are Kirchhoff's: each vortex's x and y coordinates are a conjugate pair, with the strength as a weight, and the energy is the Hamiltonian. The river surface is therefore its own phase space, in the same sense as the camel of episode 1.
The film tells the story of how the tubes arise and persist: vortices shed from the bed or from piers, smoke rings and their reconnection, and the viscous fade, a^2 plus 4 nu t. The flat picture is labelled as a model. The last act adds back the three-dimensional effects, following recent work on what dimples and scars reveal about the flow below.
How the visual was built
Two programs make the pictures. The model (vortex.js) integrates point vortices with a Scully core by the symplectic implicit midpoint rule, and advects dye particles whose enclosed area is measured. The renderer (water.js) computes the surface dip from the vortex strengths and traces rays through its Hessian: the bed brightness is 1/|det(I + k Hess eta)|, which gives a dark core with a bright caustic ring. A sun-angle offset is needed, because with an overhead sun the view ray passes through the same lens and the shadow disappears.
The film itself is a page of about 30 scenes cued to the narration, captured frame by frame with headless Chrome. The music is synthesized from scratch and ducked under the voice with sidechain compression. This draft starts and ends on Ben's own boat footage: frames from two video clips, with the dimples tracked and replayed at four times speed, grounded with swisstopo imagery and river-gauge data.
What is verified and what is not
There are no formal proofs in this episode. The simulation is checked numerically against theory: the pair speed 1.576 cm/s agrees with the formula, energy and impulses are conserved to 4e-9 and 1e-14 over 40 seconds, and the camel-shaped dye patch keeps its area to within 0.08 percent while being stretched.
The flat flow is a model, not a property of the real surface. The water renderer is an illustration. Whether the large dimples on the Rhine were shed by bed prominences is a hypothesis, supported by a published case of dimples shed from a bridge pillar on the Nidelva; the research did not locate a published dimple baseline for a river.
By the numbers
- Selected evidence: 18 sessions (2 Claude main, 12 Claude subagent, 4 Codex CLI/lane), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 24 direct human-message prompts, 4 queued typed prompts, and 3 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 1,606 logged agent output units; 1,803 tool calls; 12 new Agent requests. Output units differ between formats. Claude logged 1,237,477 output tokens.
- Models recorded: claude-opus-5-5, gpt-6-sol. Core records cover 2 active days, spanning 34.3 elapsed hours, including waits. The current spoken script has 83 lines and 780 words, counted from its text fields.
In Ben's words
“if you can make this connection legibile to an undergrad audience this would be wonderful”
2026-09-26
“I also strongly suspect the dimples we saw were not from oars”
2026-09-26
“Don't do any audio retakes let's just do script and effect stuff for now”
2026-09-27
“since you have gps that's an angle to make the video more grounded”
2026-09-27
How the AI collaboration went
This was the most research-heavy collaboration in the selected evidence. Twelve Claude subagent transcripts cover literature, visibility, site facts, geographic assets, river conditions and the recovered footage. Codex ran three substantive CLI tasks in parallel—site research, independent visual inspection and asset gathering—plus a short readiness probe. These were delegated briefs, not additional prompts typed by Ben. One Codex task independently located a dimple in the footage; another discovered that Earth Studio required an access application, so the production used swisstopo assets.
Ben's interventions repeatedly changed what the viewer would see. He asked for conservation and dimensionality to be explained from the ground up, objected to creating unlimited vortices of one sign, preferred an angled view and rejected the oar explanation for the observed dimples. He later asked the team to ground the film in the photographed place and temporarily froze audio retakes while the script and effects developed. Existing clip lengths supported a silent animatic during that pause.
Seven script JSON versions survive. The main thread contains four substantial pasted external-review deliveries and a follow-up handoff; that counts packets, not independent reviewer models. They challenged the confusion between a flat surface and two-dimensional flow, and causal claims stronger than the available river evidence. The recovered footage became the opening and closing, while the low-flow detour was reduced. The collaboration improved the argument by cutting claims and changing shots as well as adding detail.
Credits
Made by Ben Knill with Claude Code (Claude Opus) and OpenAI Codex. Boat footage: our own, 1 June 2026. Aerial imagery © swisstopo. Nidelva photograph by Klervie le Bris (Aarnes et al., JFM 2025, CC BY 4.0). Music is original and synthesized; narration is a synthetic voice.
Chapters
- 0:00 Dimples on the Rhine
- 0:43 Under the dimple
- 2:01 Smoke rings
- 3:34 The river's own smoke rings
- 4:47 What keeps one going
- 5:56 Reading the river
The Soap Computer
A soap film minimizes area, but only locally. Pins between two plates give 120-degree Steiner networks; the film finds the shortest network about half the time for six pins. Two rings hold a catenoid, a rotated catenary, until it snaps where t tanh t = 1. The catenoid root is proved in HOL Light.
Making of
The mathematical idea
A soap film minimizes area, which makes it an analogue computer for geometry, but it only finds local minima. Between two plates, pins give the Steiner problem of the shortest network joining them, with the 120-degree junctions that Plateau's laws and Taylor's 1976 theorem predict. A film can settle into a network that is a local minimum but not the shortest. For six pins on a hexagon the candidate lengths are 5, the square root of 27 and the square root of 28; the shortest is five sides of the hexagon, with no junctions at all.
Between two rings, the film is a catenoid, a catenary rotated about the axis. As the rings are pulled apart the neck narrows until the catenoid ceases to exist, at the root of t tanh t = 1. Then the film snaps to two discs.
How the visual was built
The film follows a voice-first rule: the narration (517 words, 262 seconds) fixes every shot window, so each frame is rendered once at its final length. Six shots are Blender Cycles renders at 1080p, 32 samples with GPU denoising: the daylight film, a sodium-lamp version, four pins, two six-pin rigs, the catenoid and the snap, and the drain and pop. The film is a Principled BSDF with a thin-film layer of refractive index 1.33; the thickness comes from a baked flow simulation, and a studio world that is black to the camera but bright to indirect rays lets walls at every orientation show colour. Cycles' thin film is RGB, so the sodium look is a one-wavelength Airy reflectance built from shader nodes. In all, 4,192 Blender frames took about 4.9 hours of job time, roughly 2.5 hours of wall time, on one RTX 2070 shared by two jobs.
The charts, dips and proof screens are 2D renders. The Steiner shots show simulated dips run on the bluestar machine, because node versions settle some chaotic relaxations differently.
What is verified and what is not
One result is proved in HOL Light: the catenoid snap root lies in (1.1996786402, 1.1996786403) and is the unique positive root of t tanh t = 1, with no new axioms. The Steiner optima were checked by exhaustive enumeration, with gaps below 1e-7, and the minimum-area surface code reproduces the tetrahedron value 6 sqrt 2 and the cube-square value of 0.186 of the edge, which matches Surface Evolver.
The success rates in the film are simulated dips, not a tub of soap: with six pins, 506 of the first 1,000 unshaken dips found the shortest network (50.6 percent, 95 percent confidence 47.5 to 53.7). Three, four and five pins always succeed. The thin-film colours are computed physically but not validated against photographs.
The final status qualifies the flow appearance: its flat-film shot uses procedural noise rather than a fluid simulation. The GPU soap-flow experiments and a physically computed interference material do not validate every thickness pattern in this cut.
By the numbers
- Selected evidence: 5 sessions (2 Claude main, 2 Claude subagent, 1 Codex CLI/lane), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 3 direct human-message prompts, 1 queued typed prompts, and 5 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 592 logged agent output units; 603 tool calls; 2 new Agent requests. Output units differ between formats. Claude logged 215,223 output tokens.
- Models recorded: claude-opus-5-5, gpt-6.1-sol. Core records cover 6 active days, spanning 172.4 elapsed hours, including waits. The current spoken script has 42 lines and 517 words, counted from its text fields.
In Ben's words
“surface area minimization.... catenaries? solving complex problems through free energy minimization? whatever direction you want”
2026-09-28
“the way we get the stripes instead of the rainbow in single wavelength”
2026-09-28
“we'll keep it to sims and blender for now”
2026-09-28
“we have a neat thin film effect but the variations aren't soaplike”
2026-09-29
How the AI collaboration went
Ben supplied two productive constraints early: use simulations and Blender, and show the stripes under single-wavelength light. He also noticed that the first interference material looked attractive but its variations did not look like soap. Claude developed the interactive solvers and GPU flow experiments; a Codex lane supplied a seeded success-rate grid and the catenoid proof. The grid contained fifteen cells of one thousand simulated dips each. The film used one result from it, while the remaining comparisons stayed available for inspection.
The later Opus production lane followed the voice. Its 517-word narration fixed ten shot windows before expensive rendering. Six Blender shots account for 4,192 source frames; the complete shot inventory has 6,896 frames. These are source-shot counts, not a claim that the final running time times its frame rate equals 6,896. Two jobs shared the RTX 2070, giving about 4.9 hours of summed Blender job time in about 2.5 hours elapsed.
The first speech check sent five lines back for spelling or homophone repairs. A proof shot was also rendered again with a slower digit reveal. The master passed the line-by-line speech check, but a lower-bitrate review copy introduced a “pulls”/“poles” recognition discrepancy, another reason to listen. The status also flags a small mesh-pole blemish and procedural flow in the flat-film shot. The team recorded these limits rather than treating a complete frame inventory as visual acceptance.
Credits
Made by Ben Knill with Claude Code (Claude Opus) and OpenAI Codex. Narration: synthetic voice. References: Plateau, Gergonne, J. E. Taylor (1976), and Feynman's 1983 Esalen talks and QED for the sodium-lamp idea.
Chapters
- 0:00 A film thinner than a wavelength
- 0:32 Pins between two plates
- 1:13 Six pins, dipped twice
- 1:47 How often does it find the shortest?
- 3:00 The catenoid
- 3:25 The snap and the proved root
- 4:05 Settling down is a way to compute
A Proved Swirl
The thickness of a draining, storm-stirred soap film is carried along by a given vortex flow. A corner-transport-upwind scheme with convex weights keeps an exact maximum principle, and the proved error bound for the whole 7,200-step run, 5.9e-12, sits 413,580 times below one 8-bit colour step.
Making of
The mathematical idea
The colours of a soap film encode its thickness, and in a draining, stirred film the thickness is carried by the flow like a dye. The model is the advection equation for a thickness field h in a given incompressible flow. The numerical scheme is corner transport upwind: each new cell value is a convex combination of old values, so the update obeys an exact maximum principle and the scheme cannot overshoot. A per-step error bound then composes over all 7,200 steps into a single bound, 5.9e-12, which is 413,580 times smaller than one 8-bit colour step.
How the visual was built
The reference simulation runs a proved 59-word AArch64 step in a driver, producing a 2048 by 2048 state per frame; the remote render uses an x86 build checked against those states. The film shows the thickness as soap-film interference colours in Blender Cycles. Version 1 was a 60-second 1080p film at 8 samples per pixel. Version 2, shown here, was rendered at 3840 by 2160 with 256 adaptive samples, no denoiser, from the full 2048 squared state rather than a 1024 squared preview. Against a 2048-sample reference at frame 1201 the mean error is 0.026 grey levels, compared with 0.55 for version 1 with its denoiser blotches. Rendering 1,800 frames took about five hours of wall time on the bluestar machine's RTX 2070, two jobs sharing one GPU, producing 8 GB of frames. A cross-machine check compared Blender 5.2.1 on CUDA with 5.1.1 on Metal; the interior of the film agrees to a mean absolute deviation of 0.000245 in linear light.
The narrated cut follows the voice: the simulation plays at 24 frames per second after a 3-second hold, slows evenly to a stop over the last 4 seconds, and holds under the headline. It runs 86.5 seconds.
What is verified and what is not
This is the most carefully scoped claim in the series. In HOL Light, with no new axioms, three things are proved: the machine-code step, bit for bit, together with its maximum principle and error bound; the compiled driver; and the envelope over N steps. The accompanying page lists five of six claims as proved and separates the rest as checked: all 1,801 frames reproduce bit for bit from the proved object files, on the Apple M5 and, for the frames rendered on the other machine, from an x86 build of the same C source whose state hashes are equal at every step.
What is not proved is listed in plain headings on the page and repeated here. The flow is given, not derived. This is transport, not Navier-Stokes. The distance from the numerical solution to the true solution of the PDE is not bounded. The colour map is not in HOL. Nothing outside the two functions is proved, including the x86 build that wrote the frames rendered on bluestar.
By the numbers
- Selected evidence: 5 sessions (1 Claude main, 3 Claude subagent, 1 Codex CLI/lane), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 0 direct human-message prompts, 0 queued typed prompts, and 9 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 993 logged agent output units; 1,017 tool calls; 2 new Agent requests. Output units differ between formats. Claude logged 51,083 output tokens.
- Models recorded: claude-opus-5-5, gpt-6-astra. Core records cover 3 active days, spanning 43.0 elapsed hours, including waits. The current spoken script has 14 lines and 165 words, counted from its text fields.
- Retained proof receipts: 114 files (52 passed, 61 failed, 1 incomplete); includes probes and preparation, not distinct theorem attempts.
In Ben's words
“I would very much like us to push to a program that is verified + produces even nicer visuals”
2026-10-04 · shared series/workflow direction
“what are the worst pain points as assembly programs get large?”
2026-10-04 · shared series/workflow direction
“bluestar should also take on some blender work, maybe be the first option for longer bakes”
2026-10-05 · shared series/workflow direction
How the AI collaboration went
Ben asked for a verified program that also produced better pictures. The work split into an Opus simulation-and-film lane and instruction-semantics and call-rule support. The driver proof applied the step's contract at calls rather than replaying the step inside every iteration; the recorded check counted 45 driver instructions and none from the step body. This is a concrete example of collaboration through a small contract rather than a shared claim that the whole pipeline is verified.
The second film lane tested the look before committing to a long bake. It compared sample counts and denoising against a 2,048-sample reference, selected 4K with 256 adaptive samples and no denoiser, and restored the full simulation texture resolution. Ben directed long Blender work to the remote machine. The final bake took about five hours elapsed for two jobs sharing one GPU.
An apparent frame mismatch then exposed a faulty check. Near frame 60, adjacent pictures differed less than the first film's compression error. The agent changed the check to align fifteen-frame stretches and tested it with both a harmless re-encode and a deliberately dropped frame. One passed and the other failed. That repair checked whether the test could distinguish an actual sequencing error. The narration was also tightened to say that the remote x86 states were checked equal to the proved Arm computation, preserving the boundary between theorem and cross-machine experiment.
Credits
Made by Ben Knill with Claude Code (Claude Opus) and OpenAI Codex. Renders: Blender Cycles. Proofs: HOL Light with Hearth. Narration: synthetic voice.
How wrong can a dot product be?
The same dot product summed in different orders rounds differently. Ubuntu's shipped Arm64 DDOT, with 16 FMA lanes and a tree, has a proved worst-case error bound, and constructed inputs reach 99.635 percent of it at n = 4096.
Making of
The mathematical idea
Floating-point addition is not associative, so a dot product computed in a different order rounds differently. The classical analysis gives a worst-case bound of the form gamma_n times the sum of |x_i y_i|. The kernel here, the tested Arm64 DDOT from Ubuntu's OpenBLAS, sums with 16 fused multiply-add lanes and a tree, and its own error bound can be proved for exactly that order. The film then asks how close real inputs can get to the bound.
Constructed inputs reach 99.6351252737 percent of the full bound at n = 4096. The film is explicit that this probes underflow and the shape of this kernel, not the optimality of the classical gamma bound.
How the visual was built
The film is a 210-second, 720p, 24 frames per second render of 5,040 frames from 2D graphics, cut to a narration of 392 words that passed a local speech-recognition check on all 33 clips. A companion page lets the viewer drag inputs and see the rounding change, and its 102 arithmetic cases are compared with an independent exact oracle.
What is verified and what is not
In HOL Light, through Hearth with the s2n-arm-fp64 profile and no new axioms, the DDOT error theorem is proved for the kernel's machine code, and a witness file evaluates the n = 1 and n = 4096 witnesses inside HOL. The 51 witnesses (best observed input at 17 lengths) also reproduce bit for bit under qemu and on the Apple M5, which ties the proved model to a real machine.
The closing section reports a different kind of result. Three identical all-ones calls to sdot returned 64, 128 and 192 on the native M5, because the shipped sdot kernel adds a leftover register. This was measured, not proved, at the time of the film; the companion film on this page, Proved vs Shipped, follows up with the proved statements. Three shipped kernels were affected in the campaign, and 53 other dot variants were not.
By the numbers
- Selected evidence: 5 sessions (2 Claude main, 1 Claude subagent, 2 Codex CLI/lane), including named support and shared coordination. This is a conservative attributed set, not all historical work.
- 0 direct human-message prompts, 0 queued typed prompts, and 5 shared directions. Delegated briefs and automatic notifications are excluded from these counts.
- 551 logged agent output units; 507 tool calls; 0 new Agent requests. Output units differ between formats. Claude logged 10,730 output tokens.
- Models recorded: claude-opus-5-5, gpt-6-astra, gpt-6.1-sol. Core records cover 2 active days, spanning 12.4 elapsed hours, including waits. The current spoken script has 33 lines and 392 words, counted from its text fields.
In Ben's words
“I'm interested in your thoughts on verified assembly simulators for visualizing interesting dynamics.”
2026-09-25 · shared series/workflow direction
“the idea is to have a few nice examples, and maybe some live development, and a description of the principles of development in this area.”
2026-10-05 · shared series/workflow direction
“I wanna do a 99R push to generically improve the presentations/videos and get everything in one place for me to review.”
2026-10-05 · shared series/workflow direction
How the AI collaboration went
The mathematical lane and narration desk had different jobs. Codex developed the witnesses, exact comparisons, page and provisional film on the remote machine. The Claude desk handled the local series voice and the spoken-number edit, then sent audio back for composition. Its brief also included a Talbot film; only dot-product-related records are counted here. A shared desk is not evidence that every action in its transcript belonged to this episode.
The witness file turned the headline examples into HOL-evaluated results, while native and emulated replay connected the abstract arithmetic model to observed bits. The interactive page supplied another independent check: 102 browser cases against an exact oracle. The sdot surprise arrived from a neighboring lane and was added as act three. At that stage its repeated-call behavior was a measurement. The later sdot theorem is not retroactively presented as evidence available to this film's first cut.
The 654-word initial narration source was edited to a current 392-word spoken script in 33 clips. Spoken numerals and abbreviations needed attention even though the rendering was mostly two-dimensional graphics. The final cut had 5,040 frames and ran 210 seconds. Ben's recoverable prompts here mainly concern the broader aim, course fit and review workflow. The sampled series directions are useful context, but they do not establish that he personally selected the particular adversarial inputs or authored the bound.
Credits
Made by Ben Knill with Claude Code (Claude Opus). Narration: synthetic voice. The kernel analysed is OpenBLAS's Arm64 DDOT as shipped by Ubuntu.
Chapters
- 0:00 A dot product looks simple
- 0:30 Different paths, different roundings
- 1:10 The proved bound
- 1:45 Inputs that nearly reach it
- 2:25 How the proof was made
- 2:55 The sdot surprise
- 3:20 Close