Proposal video, behind the scenes
This video was made by an agent I built.
It reads the job post, writes the script in my words, voices it with a clone of my voice and lip-syncs my face. I check it at two gates. Every note I give it becomes a rule it keeps.
74of my real sales videos it learned my delivery from
204corrections it has turned into rules
145automated tests run on every change
~$2to make one video, voice and lip-sync included
The pipeline
From a job post to a finished video
Read the job
The post comes from Upwork's official API. The agent writes down their situation in their words, what's broken, and three possible insights, then picks one.Pull the proof
Case studies come from my knowledge base for every video, never from memory. Every number in the script has to trace back to a note.Write the script
My fixed lines (hello, background, close) stay word for word. The agent writes only the job-specific parts, in the rhythm of my real videos.Lint it
A checker runs every rule I've given it. An error blocks the voice, so a bad script can't cost a take.Second opinion
A separate reviewer agent reads it cold, the way the client will, and flags claims the post doesn't support.Voice
One continuous take from a professional clone of my voice. The fixed lines come from a bank of real takes, joined at the screen changes.My first approval gate
I get the script, the diagram, the audio and the cover letter in one message.Lip-sync and render
My face is lip-synced on ten GPUs in parallel, then laid over the job post, the diagram and my old company's site as a webcam bubble.I watch it and send it gate
Nothing reaches a client without me. After I send it, the video goes into a test that tracks AI videos against my real ones.
The agent folder
What it reads before it writes anything
The context lives in files, not in a prompt. Each kind of rule has exactly one home, so a correction never has to be made twice.
vibecast-agent/
├── .claude/skills/video-script/
│ ├── SKILL.md # how a script gets written, step by step
│ ├── playbook.md # my real lines and habits, from 74 of my videos
│ ├── blocks.txt # my fixed lines, word for word
│ ├── contracts.json # how long each part is and how it opens and ends
│ ├── lint-rules.json # everything a machine can check
│ └── lessons.md # every correction I've given, verbatim, and where it went
├── .claude/agents/ # reviewer, diagram designer, cover-letter writer, job triager
├── scripts/<job>/ # one folder per job: the read, the proof, the script, the letter
├── scripts/approved/ # videos I approved or rejected, re-checked on every test run
└── vibecast-engine/ # voice, lip-sync and render (TypeScript, GPUs on Modal)
+ the company brain: my knowledge base, 426 notes of case studies, clients and calls
The correction loop
When a run is only 70% right
I don't just tweak the prompt. A note goes through the same five steps every time, and the test is what keeps the mistake from coming back.
1
Save my words
Verbatim, with the date and the video.2
Name the class
The principle behind the note, so it covers wordings I never used.3
One home
A fixed line, a size limit, a lint check or a playbook move.4
A test
The check gets tested on new wordings before it ships.5
Re-check the library
Every approved video must still pass. Every rejected one must still fail.From the ledger
A few of the 204
L-192
That's a typical 3-2, AI pattern… We should have a rule preventing that.Tidy lists of three are now a lint error anywhere I speak.
L-184
Never quote price in proposals or cover letters.Any currency in a letter fails its own checker.
L-151
You should say something like, 'And one thing I was wondering'… that transition is pretty not the smoothest.A question to the client needs a lead-in, or lint stops the voice.
L-203
Blind B is slightly better.Speech marks (a hang on a word, one stressed word, a quick "right?") won a blind listen against my approved take, so the writer uses them now.
This video's own lint report
The check it passed before the voice
2026-10-09_marketing-engineer-growth: 288 words · ~89 s · 5 paragraphs · 4 fillers
warn L-105 5 paragraphs; the default path is 4
warn L-011 24-word sentence; the clone goes flat on long ones
OK with warnings
Experiments, with kill rules set first
Nothing changes without a fair test
Voice providersThree challengers trained on the same 55 minutes of my voice, judged blind. The current one kept its seat.
Speech marksThe same script with and without, judged blind. The marks won and went into the writer.
GPU releaseStopping each GPU after its own chunk cut lip-sync cost 16% with no visible change, checked frame by frame.
AI vs real videosEvery AI proposal is logged against 122 real ones I sent this summer (40% viewed, 19% replied).