← Math in Sight / A Proved Swirl

A studio diary · October 10, 2026 · All times EDT

Finding a voice
by listening.

A human ear, an AI experiment builder, and a live voice.
What changed when we stopped giving so much direction?

Interference colours in the soap film from A Proved Swirl
The picture was already there. This afternoon was about how it should sound.

2 h 49 minFirst sampler request → final rating

104 local samplesFour rounds of 26

4.8 / 5Ben’s rating of this Sol take

At 11:36 a.m., the brief was practical: find a small, fast local voice and make it easy to audition. By 2:25 p.m., the best result was a full-script live performance with almost no direction. The route between them was a succession of listening tests, blunt reactions, and changes to what the next test should measure.

The AI could produce candidates and assemble a film. Ben’s feedback determined which differences mattered: clarity versus realism, a room that sounded attached to the speaker, and the continuity of an actual performance.

The selected take

Sol, with room to speak.

86.5 seconds. The 1080p Sol cut is served by YouTube. Compare the original YouTube narration ↗

“We’d like you to perform this script for us.
Unhurried documentary narration.”

The complete style direction. The full script followed in the same typed turn.

The afternoon, in six moves

The ear kept changing the experiment.

01

Make the audition blind.

“actually have 26 of them and the labels are letters of the alphabet only.”

Ben → Claude, 11:36 a.m.

Claude built a local HTML listening page with A–Z samples, automatic advance, and copyable ratings. The first rounds mixed Kokoro, Supertonic, Pocket and Soprano voices with cheap processing. The next brief replaced integer stars with a continuous 0–5 slider. A later comment box let Ben say what a number could not.

The real local audition interface, with A–Z buttons, playback, a rating slider and a comment box
The round-four interface, captured again for this diary. Letter-only labels hid the voice and processing recipe; saved comments supplied the interpretation.

The first round’s “humanize” chain averaged 4.25. In round two, raw samples averaged 4.03 and humanize 3.66. The samples differed across rounds, so this is no controlled head-to-head effect estimate. It was enough reason to question whether more processing was helping.

02

Turn an adjective into a technical change.

“the background noise is not part of the spoken performance on the soundstage idk how to say it”

Ben’s round-three comment, 12:34 p.m.

“More human” became a specific mixing problem. Room tone added after the voice chain sounded detached. The next round put room sound through the microphone chain with the voice, narrowed the search to af_heart and bm_lewis, and compared lavalier, condenser and close-mic colouring.

Ben liked the dry lavalier’s clarity and the condenser’s timbre on af_heart; bm_lewis could sound like a voice actor. Inserted pauses sounded “weird,” and a de-esser made little audible difference. The useful output was a few reusable presets, not a maximal effects stack.

Same local voice, dry lavalier

Kokoro af_heart · sample G, round 4.
Ben: “very good … very clear.”

Same line, condenser + room

Kokoro af_heart · sample I, round 4.
Ben liked the added realism.

03

Let the live voice warm up.

The work moved to GPT-Live through the Codex live-voice connection. An elaborate “voice actress” prompt produced performances Ben found fake. At 1:17 p.m., he replaced it with a simple request:

“we’d like you to perform this line for us: ‘line’”

Ben’s suggested wording, 1:17 p.m.

Record the whole reply, allow its “Absolutely, Ben…” preamble, then cut to the requested line. Preserve the pauses inside it: “dont cut out pauses btw.” The second short-line batch had twelve takes across nine voices. Ben called the overall quality very high; natural short lines still did not settle how a whole film should sound.

04

Exact words were not enough.

The first full-film experiment used a local male voice to speak the instructions. The quotation boundary disappeared in speech, and the live voice improvised. Typed turns fixed that: eight separate Vale replies matched the requested paragraphs in the session transcript and could be placed at the original beat starts.

“revoice is very bad. try directing them w/ text again and to the full script and one take.”

Ben, 1:45 p.m.

A text match answered the engineering question. It did not answer the listening question. Ben changed the unit of performance from a paragraph to the whole script.

The rejected approach

12 seconds from the typed, separate-beat Vale assembly. Different voice and delivery; this is an illustration of the experiment, not a controlled voice comparison.

The selected approach

12 seconds from Sol’s continuous take. Preamble trimmed; internal timing left intact.

05

Less direction. More continuity.

“continuous is always better.”
“we are overprompting”

Ben, 2:00 p.m. · judgments within this session

Claude compared continuous one-take readings with versions split at paragraph pauses and shifted back onto the old beat starts. Ben preferred continuous delivery; Ember was the best take in that round. The next batch used only “Unhurried documentary narration” plus the script.

Sol4.8 of 54.8
Maple3.7 of 53.7
Ember1.4 of 51.4

Arbor was unscored: it changed the opening and mispronounced “H-O-L Light.” These are Ben’s ratings of these takes, not a benchmark or stable ranking of voices.

Sol won: “best so far,” followed by “reeealy good.” Maple’s voice changed over the narration; Ember fell from the prior round’s favourite to 1.4. Recording several takes and listening remained necessary.

The actual round-three comparison page with Ember, Sol, Maple and Arbor film players
The comparison interface recreated for this diary, with film posters and Ben’s saved October 10 ratings restored.
06

Keep the failed idea in the story.

In a parallel experiment, the live voices listened to the best local male model read the script, then received “Done. Now your own take:”. Ember and Maple treated “take” as an invitation to give an opinion. Maple even chatted during playback. The longer listen-then-perform setup failed as run.

That does not establish that audio references cannot help. The wording was ambiguous. The saved notes propose “Now please perform it yourself, in your own style:” for a future attempt; that repair was not tested in this session.

What the cooperation looked like

The feedback was part of the machinery.

Ben listenedNamed the audible problem.
Rejected and selected takes.
Claude builtMade the next comparison.
Recorded, cut and assembled.
GPT-Live performedProduced candidates through
the Codex voice connection.

The decisive information was often a few informal words: room noise felt separate, pauses felt robotic, the prompt felt overbearing. Each observation became a change to the next comparison. The human was specifying the target by reacting to something audible; the agent made those reactions cheap to test.

One delegation researched microphone processing and another preserved the lab’s tools and results. The core listening loop stayed in the main Claude session. Codex supplied the live-voice connection; this diary was assembled afterward by Codex from the session record, saved ratings, and actual media.

The working recipe: type the whole script in one turn, give a short style cue, record the full reply, trim its preamble, preserve the performance’s internal pauses, and audition several takes. Check words and pronunciation, then listen to the film.

Sources, timing and limits

The diary uses the October 10 tail of Ben’s long Claude Code session, the four local sampler keys and result files, the two live-voice batches, and the Swirl round-three saved ratings. Four local rounds contained 26 samples each; the short live rounds contained 11 and 12 takes. Some later rounds have comments rather than numeric ratings.

The 2 h 49 min span runs from the 11:36 a.m. sampler request to the 2:25 p.m. final rating; it includes listening, generation, waits and unrelated work in the same session. It is not active labour or model runtime. The sampler screenshot shows the real local interface reopened for this diary. The comparison screenshot restores saved ratings and adds film posters for readability. Neither was captured during the original auditions.

The Sol session transcript matches the supplied wording after its preamble, apart from typography. A transcript is not an independent audit of pronunciation. This page preserves Ben’s judgment of the heard performance. The film’s existing proof claims were not rerun for this audio change.

Read the selected prompt excerpts and timestamps. The complete private conversation is not published.