Case study · Posterity
Building an AI Editorial Desk: What I Automated, What I Kept Human, and What Broke
How a small multi-agent newsroom researches, writes and produces history Shorts while a human still owns the taste.
Muhammad Gane9 min read
Posterity is a faceless short-form history channel. It tells cinematic stories about dark history, vanished kingdoms, mythology and folklore, documented where the record is solid and labelled where it's legend. Its line is History remembers the victors. We remember the rest.
This write-up isn't really about the videos. It's about the system that makes them: a small AI editorial desk of specialist agents, one Editor-in-Chief that owns the quality bar, and me, sitting at a few deliberate gates. It has been publishing since September 13, 2026. Some of the design held up and some of it didn't. The parts that didn't are the more useful half of this article.
Project record
Posterity
History remembers the victors. We remember the rest.
Cinematic stories told in under ninety seconds, documented where the record is solid and labelled where it's legend.
- Subject
- Dark history, vanished kingdoms, mythology, folklore
- Format
- Vertical 9:16 Shorts, cross-posted to TikTok and Reels
- Cadence
- Most mornings, when an episode clears the gate
- Operated by
- Five agents and one human
The problem with one agent
A good 75-second history Short takes a lot more than a script. Someone has to find a story worth telling and check that it can hook a viewer in under two seconds. Then it has to be researched, and someone has to think about which imagery can legitimately be used. After that comes the voiceover, a shot-by-shot visual plan, the images themselves, narration, the edit, a quality check, and publishing to three platforms with different rules. Then someone reads the results and uses them to choose the next story.
The obvious design is to hand one capable agent the whole job. I didn't, and the reason has less to do with capability than with inspection. When one agent owns everything, a weak episode has no address. Was the topic wrong, the research thin, the script flat or the edit sloppy? Every correction lands on the whole thing at once, and the agent's context fills with every stage's details whether or not they matter to the step it's on.
So I split the work the way a newsroom would, by stage and by ownership.
The desk
Five agents, one human. Each specialist owns a stage:
- Archivist finds and scores candidate stories, builds fact packs and thinks about rights.
- Scriptor writes the voiceover, hooks, titles and sequence briefs.
- Atelier assembles the cut, runs QA, produces masters and handles publishing operations.
- Pulse keeps the metrics ledger, forms growth hypotheses and scans comparable channels.
Above them sits Vellum, the Editor-in-Chief and showrunner. Vellum owns editorial direction and the quality bar, greenlights topics, and gates cuts before I see them.
Owner · human
Muhammad Gane
Sets the bar, makes the plates, clears the gates
Editor-in-Chief · showrunner
Vellum
Editorial direction, quality bar, owner thread, publish gate
Posterity Desk · specialists
Research
Archivist
Discovery, scoring, fact packs, rights
Writing
Scriptor
VO scripts, hooks, titles, sequence briefs
Production
Atelier
Assemble, QA, masters, publish ops
Growth & analytics
Pulse
Metrics ledger, hypotheses, competitive scan
Internal coordination. No specialist messages the owner directly.
The most important rule is about communication, not capability. Specialists can talk to each other as much as they need to, but only Vellum talks to me day to day.
That sounds like etiquette. In practice it's what keeps the desk usable. Without it, five agents each report status, each ask small questions, and each raise decisions that belong together in separate threads. I end up as the integration layer, reconciling fragments across chats. With it, there is one thread and one voice accountable for the state of the desk, and decisions reach me already framed. It also saves tokens, because parallel status updates are pure overhead.
The packet is the interface
Every morning the desk delivers an INFO PACKET. It looks like a status update, but it works more like an executable spec for one episode:
- why this story, and why now
- the narrator's arc
- the full voiceover
- five to seven sequences, each with a brief for the imagery I'll generate, tagged as evidence, reconstruction or map
- bans: what the episode must avoid
- thumbnail concepts, metadata and a music note
It arrives alongside the final script, and both come as complete documents. A summary doesn't count, because I'm reviewing the actual words that will be spoken and the actual briefs I'll be working from.
INFO-PACKET.md
Spec
- 01Why this story now
- 02Narrator arc
- 03Full voiceover
~65–90 seconds
- 04Sequences
5–7, each with a generation brief
EvidenceReconstructionMap
- 05Bans
- 06Thumbnail concepts
- 07Metadata
- 08Music direction
script-final.md
Voiceover
Why a document
Every downstream step reads from it: my image generation, the assembly, the mute gate and the metadata. If something is wrong, it is wrong in one place I can point at.This is the most transferable thing Posterity has taught me about agent systems: a well-structured intermediate artifact is a better coordination surface than conversation. When agents talk freely to each other, they drift, repeat themselves and lose decisions in the scrollback. A document with fixed sections can be reviewed, compared with yesterday's, handed to the next stage intact, and corrected in exactly the right place. If a thumbnail concept is wrong, I know which section to fix and who wrote it.
Where I stay in the loop
The desk is deliberately not autonomous. Today's pattern:
- The agents research and prepare the packet.
- I review it.
- I generate and select the visual plates, including a unique outro.
- Atelier assembles the rough cut.
- Vellum runs the mute gate.
- I take a final look.
- Only then does it publish.
Automation is the desk. Taste is the gate.
Plates show most clearly where that line sits. Early on, the agents produced pools of candidate images on contact sheets for me to choose from. Since September 20, I make the plates myself. The packet gives me a brief for each sequence, I draft the imagery with dedicated image models, and I hand the results back for the desk to fold into the story. Thumbnails went the same way, and I now design and upload them by hand. Neither change was a statement of principle. Both steps moved to me when the agent output stopped meeting the bar, which I come back to below.
How a Short gets made
An episode moves through five phases. Agents do most of the steps. The gates are where I, or Vellum on my behalf, can stop it.
IPrepare
- Agent
Discover & score
Shortlist stories that can teach something in the first two seconds
Archivist · Vellum greenlights
- Agent
INFO PACKET + VO
Research depth, full voiceover and a brief for every sequence
Scriptor + Archivist
IIMake
- HumanGate
Plates
I generate and select the stills, plus a unique outro
Owner
- Agent
Assemble
VO, spoken captions, music and outro into a rough cut
Atelier
IIIGate
- AgentGate
Mute gate
Does the cut teach with the sound off? Failures stay internal
Vellum
- HumanGate
Final look
Nothing goes public until I've watched it
Owner
IVPublish
- Platform
YouTube first
Short goes live on the channel with a custom thumbnail
Owner · Atelier
- Platform
TikTok + Instagram
Only after YouTube is public; Instagram forced to 9:16
Atelier
VLearn
- Agent
Ledger → next story
Metrics and the experiment log feed the next pick
Pulse → Vellum
Prepare
- Agent
Discover & score
Shortlist stories that can teach something in the first two seconds
Archivist · Vellum greenlights
- Agent
INFO PACKET + VO
Research depth, full voiceover and a brief for every sequence
Scriptor + Archivist
Make
- HumanGate
Plates
I generate and select the stills, plus a unique outro
Owner
- Agent
Assemble
VO, spoken captions, music and outro into a rough cut
Atelier
Gate
- AgentGate
Mute gate
Does the cut teach with the sound off? Failures stay internal
Vellum
- HumanGate
Final look
Nothing goes public until I've watched it
Owner
Publish
- Platform
YouTube first
Short goes live on the channel with a custom thumbnail
Owner · Atelier
- Platform
TikTok + Instagram
Only after YouTube is public; Instagram forced to 9:16
Atelier
Learn
- Agent
Ledger → next story
Metrics and the experiment log feed the next pick
Pulse → Vellum
- Back to discovery with the next story
The publishing order is a standing rule, not a preference. YouTube goes first. TikTok and Instagram follow only once the YouTube Short is public, never alongside a draft. Every upload gets an open discussion question pinned under it. YouTube Shorts is the channel's growth priority, so everything else waits for it.
The Recovered Chronicle
The visual system has a name, The Recovered Chronicle, which gives every stage the same reference point. It has three layers, each with a job, and a short list of standing rules that turn parts of taste into something checkable.
I
Evidence
Real archaeology, coins, maps and period sources, where the rights allow it.
II
Reconstruction
Painterly, tactile generated stills for the shots the historical record can't supply.
III
Editorial graphics
Maps, labels and timelines that add information instead of repeating the narration.
Standing rules
- Mute test
- The cut has to teach something with the sound off.
- Full spoken captions
- Captions carry every word of the VO, not just place names.
- Unique outro
- Every episode gets its own closing plate. No shared brand card.
- AI disclosure
- Platform labels and descriptions disclose generated imagery where required.
- True 9:16 on Instagram
- Force the crop before sharing; the default is square.
- Skip, don't ship weak
- A missed day costs less than a weak episode.
- Pin a question
- Every publish gets an open discussion question pinned.
None of this automates taste. What it does is turn the parts of taste that can be written down into rules any stage can check. That leaves my attention for the parts that can't be written down: whether a plate feels right, whether the story lands, whether the ending earns its question.
The feedback loop
Each morning, before the next packet is written, the desk runs a scheduled loop over analytics and the competitive landscape. Pulse updates the ledger, scans comparable channels and turns what it finds into a hypothesis. Vellum then decides whether that hypothesis fits the channel before it shapes the next pick.
- 01Desk + ownerPublish, then pin a discussion question
- 02PulseLogs the experiment in the metrics ledger
- 03PulseScans comparable channels for what's moving
- 04PulseTurns both into a growth hypothesis
- 05VellumWeighs it against the brand lane and picks
The pick becomes tomorrow's packet
Now · exploring
Early in the channel's life, Vellum tries new angles and watches what holds. One quiet day doesn't rewrite the thesis.
Later · a formal series
With weeks of accumulated data, the picks should settle into recurring series instead of one-off experiments.
The key design choice is where the brand sits in that loop. It constrains the output. The metrics don't get to rewrite it. That distinction was tested within the first week.
Watch the desk's output
These are the episodes shipped so far, in publishing order. The lane labels show the arc: two Roman myth-busts early, then a deliberate turn toward forgotten kingdoms and places told as stories.
Start here · No. 06 · Sep 21, 2026
Before Alexandria, This Port Ruled Egypt's Mouth
Plays here. Nothing loads from YouTube until you press play, and only one player exists at a time.
In publishing order
Myth-bustStory-led
No. 01 ·
He Found Troy by Destroying It
evergreen ironyNo. 02 ·
Carthage — He Never Salted It
myth-bustNo. 03 ·
Nero — He Never Fiddled
Roman myth-bustNo. 04 ·
Great Zimbabwe
forgotten kingdomNo. 05 ·
The Fourth Power History Mislaid
forgotten kingdomNo. 06 ·
Before Alexandria, This Port Ruled Egypt's Mouth
forgotten portNo. 07 ·
This Capital Had No Roads
forgotten capitalNo. 08 ·
Chan Chan: The Mud Capital the Empire Took
adobe capital / conquest
What I learned running it
One thread to the human is worth more than it sounds
With five agents, the cost that grows fastest isn't compute. It's my attention. Routing everything through Vellum means decisions arrive grouped, framed and already argued. The work stays fast without my week disappearing into status updates.
Rule nowSpecialists coordinate with each other. Only the Editor-in-Chief brings things to me.
The human gates didn't disappear. Some came back.
The desk genuinely sped up research, scripting and assembly. It didn't remove the need for judgment, and in two places it handed work back to me: plates moved from agent-made contact sheets to images I generate from the packet's briefs, and thumbnails became a manual job.
I also set up a temporary early window in which I personally review every upload before it goes public, with a plan to reconsider after the fifth. I'm no longer sure I'll drop it.
Rule nowWhen agent output keeps missing the bar, that step moves to a human gate.
Early packaging signals can mislead
Roman myth-busting converted well early. Carthage and Nero, both "he never actually did that" stories, got real early traction, and the tempting move was to pivot to more myth-busts every day. I steered the desk back instead. Posterity's lasting lane is dark history and forgotten kingdoms told as stories, not an endless run of corrections. Aksum, Thonis-Heracleion, Nan Madol and Chan Chan are that lane.
A related lesson: a brand-new channel does get some distribution in the Shorts feed, just unevenly. One quiet day isn't evidence that the thesis is wrong, and patience beats rewriting everything after a slow upload.
Rule nowMetrics inform the next pick. They don't get to redefine the channel.
One master file doesn't mean one publishing behavior
The same vertical master goes to all three platforms, which suggests the publishing step is the same job done three times. It isn't. Instagram's create flow defaults to a square crop and letterboxes a 9:16 video unless you change it, so forcing 9:16 before sharing became a hard rule. Early TikTok posts often hit a cold start with slow indexing, which is easy to misread as "history doesn't work on TikTok." And cross-posting waits until the YouTube Short is live, because the order is part of the strategy.
Rule nowEach platform gets its own checklist, even when the file is identical.
Silence is a QA feature
Vellum mute-gates every rough cut: with the sound off, does it still teach something? When a cut fails, Vellum doesn't tell me. It goes back to the desk, and I only hear about cuts that passed. That can look like withholding information, but it actually removes noise. If every internal attempt reached me, I'd be back to being the QA department, and each cut I saw would carry less signal.
Rule nowEscalate approved work, not every attempt.
Quality drifts as the context grows
This is the failure I understand least and notice most. Over time, the agents get noticeably sloppier, and in the same places. Thumbnail concepts get weaker, generated images lose the art direction, video and imagery drift off the phone's 9:16 framing, and the rough-cut gate starts letting through things it caught a week earlier. The longer a thread's context grows, the more human intervention the same work needs.
I can't reset sessions in the agent harness I use, and I suspect that's part of the cause. Standing instructions from early in a long context seem to lose weight against everything piled on since. What I do for now is step in more often, keep standing rules in written files the agents re-read instead of relying on conversational memory, and keep my own look before every upload. None of that is a fix. It contains the problem while I work out what actually causes it.
It's also the main reason I'm cautious about long-form. If a 75-second Short drifts over a few weeks, I don't yet know what happens across an hour-long story.
Keeping a human in the loop has a bill
Keeping myself in the loop is the right call for quality, but it isn't free. Every correction, re-brief and round of back-and-forth costs tokens, on top of the agents' own inefficiency, and I regularly hit my weekly usage limit around day five. The system works. It's also nothing like the cheap, passive machine the "start a faceless channel and let AI print money" pitch describes.
Rule nowBudget for the human loop, in hours and in tokens.
What this says about agent design
Posterity is a media project, but most of what it taught me applies to any agentic system that does real work:
- Design who talks to whom, not just the agents. Who is allowed to talk to the human matters as much as what each agent can do.
- Coordinate through artifacts. A fixed-shape document can be inspected, reviewed and corrected in a way a conversation can't.
- Place gates on purpose. Put human judgment where taste and risk concentrate, and let agents gate each other quietly everywhere else.
- Write taste down where you can. Named systems and standing rules let every stage check the parts of quality that can be checked.
- Expect degradation. Long-running agents don't hold their quality by default. Plan for re-grounding and extra review from the start.
- Count the human's cost. A human-in-the-loop design only lasts if the loop is affordable.
A system can be meaningfully agentic without pretending the human has left the room. Posterity works because I stopped trying to remove myself and started deciding exactly where I belong in it.
What's next
Two open questions. The first is content selection. The desk is still exploring, trying angles and watching what holds. As weeks of analytics accumulate, I expect the picks to settle into a formal set of recurring series.
The second is long-form. I'd like to take the same kind of story to an hour. Whether this desk can hold its quality over that length, given what I've seen of drift, is something I'll find out one day at a time.