Proposal video, behind the scenes

This video was made by an agent I built.

It reads the job post, writes the script in my words, voices it with a clone of my voice and lip-syncs my face. I check it at two gates. Every note I give it becomes a rule it keeps.

74of my real sales videos it learned my delivery from
204corrections it has turned into rules
145automated tests run on every change
~$2to make one video, voice and lip-sync included

The pipeline

From a job post to a finished video

  1. Read the job

    The post comes from Upwork's official API. The agent writes down their situation in their words, what's broken, and three possible insights, then picks one.
  2. Pull the proof

    Case studies come from my knowledge base for every video, never from memory. Every number in the script has to trace back to a note.
  3. Write the script

    My fixed lines (hello, background, close) stay word for word. The agent writes only the job-specific parts, in the rhythm of my real videos.
  4. Lint it

    A checker runs every rule I've given it. An error blocks the voice, so a bad script can't cost a take.
  5. Second opinion

    A separate reviewer agent reads it cold, the way the client will, and flags claims the post doesn't support.
  6. Voice

    One continuous take from a professional clone of my voice. The fixed lines come from a bank of real takes, joined at the screen changes.
  7. My first approval gate

    I get the script, the diagram, the audio and the cover letter in one message.
  8. Lip-sync and render

    My face is lip-synced on ten GPUs in parallel, then laid over the job post, the diagram and my old company's site as a webcam bubble.
  9. I watch it and send it gate

    Nothing reaches a client without me. After I send it, the video goes into a test that tracks AI videos against my real ones.

The agent folder

What it reads before it writes anything

The context lives in files, not in a prompt. Each kind of rule has exactly one home, so a correction never has to be made twice.

vibecast-agent/ ├── .claude/skills/video-script/ │ ├── SKILL.md # how a script gets written, step by step │ ├── playbook.md # my real lines and habits, from 74 of my videos │ ├── blocks.txt # my fixed lines, word for word │ ├── contracts.json # how long each part is and how it opens and ends │ ├── lint-rules.json # everything a machine can check │ └── lessons.md # every correction I've given, verbatim, and where it went ├── .claude/agents/ # reviewer, diagram designer, cover-letter writer, job triager ├── scripts/<job>/ # one folder per job: the read, the proof, the script, the letter ├── scripts/approved/ # videos I approved or rejected, re-checked on every test run └── vibecast-engine/ # voice, lip-sync and render (TypeScript, GPUs on Modal) + the company brain: my knowledge base, 426 notes of case studies, clients and calls

The correction loop

When a run is only 70% right

I don't just tweak the prompt. A note goes through the same five steps every time, and the test is what keeps the mistake from coming back.

1

Save my words

Verbatim, with the date and the video.
2

Name the class

The principle behind the note, so it covers wordings I never used.
3

One home

A fixed line, a size limit, a lint check or a playbook move.
4

A test

The check gets tested on new wordings before it ships.
5

Re-check the library

Every approved video must still pass. Every rejected one must still fail.

From the ledger

A few of the 204

L-192
That's a typical 3-2, AI pattern… We should have a rule preventing that.Tidy lists of three are now a lint error anywhere I speak.
L-184
Never quote price in proposals or cover letters.Any currency in a letter fails its own checker.
L-151
You should say something like, 'And one thing I was wondering'… that transition is pretty not the smoothest.A question to the client needs a lead-in, or lint stops the voice.
L-203
Blind B is slightly better.Speech marks (a hang on a word, one stressed word, a quick "right?") won a blind listen against my approved take, so the writer uses them now.

This video's own lint report

The check it passed before the voice

2026-10-09_marketing-engineer-growth: 288 words · ~89 s · 5 paragraphs · 4 fillers warn L-105 5 paragraphs; the default path is 4 warn L-011 24-word sentence; the clone goes flat on long ones OK with warnings

Experiments, with kill rules set first

Nothing changes without a fair test

Voice providersThree challengers trained on the same 55 minutes of my voice, judged blind. The current one kept its seat.
Speech marksThe same script with and without, judged blind. The marks won and went into the writer.
GPU releaseStopping each GPU after its own chunk cut lip-sync cost 16% with no visible change, checked frame by frame.
AI vs real videosEvery AI proposal is logged against 122 real ones I sent this summer (40% viewed, 19% replied).