White Paper · Scout Fleet Deep Dive

The Forge
Scout

Media intelligence that cites its sources, frame by frame.

This is a companion to Beyond the Agent, which argued that the Scout (persistent, collective, self-improving, accountable) is the unit that scales agentic AI. Here we take the Scout pointed at the messiest input there is: recorded media.

The Forge Scout is that idea pointed at video, audio, documents, images, and 3D models. Its job is to turn hours of recording into a short, readable brief you can act on, where every line points back to the moment in the source it came from.

Another Scout in the same fleet points at innovation instead of media: the Lighthouse Scout works both sides of a match, helping a corporation find the startup that fits its problem and a startup find the corporate that genuinely needs what it builds, and it shows the real fit or says plainly there is none.

Here is the simplest way to picture the problem.

Hand an AI tool a two-hour recording and it will give you a tidy summary. It reads well. It is also a black box: you cannot tell which lines come from something actually said on the recording, which were softened or invented to make the prose flow, and which moments it never really watched. To trust it, you have to sit through the whole thing yourself. So the summary saves you nothing on the only part that matters, which is being able to point to where each claim came from.

A Forge Scout works the way a careful producer works. On a video it listens for a timed transcript, reads the text on screen, and samples frames across the whole runtime, so it has watched the thing rather than skimmed it. A media engine then lays out a grounded list of the moments that could matter, the Scout's governed model picks the ones that do and writes them up in plain language, and every line it writes is tied to a timecode in the source. You can click back to the exact moment and check it.

And when the engine cannot pull anything usable out of an asset, the Forge Scout does not paper over it. It says so, plainly, rather than inventing a moment that was never there.

A tool asks you to trust its summary. A Forge Scout shows you the proof.

00 / Executive summary

The trust problem in media summaries

Reading a recording is not a writing task. It is a sourcing task. The output is something a reader will repeat or act on, and the value is entirely in whether each line can be traced back to what was actually shown or said.

That is exactly where a lone AI tool fails quietly. It is fluent, so it produces something that looks like a faithful recap. But fluency is not faithfulness. One model, working alone over a long video, has every reason to fill a thin patch with a plausible sentence rather than admit it lost the thread. You get a recap you cannot check, which means a recap you cannot stand behind.

The Forge Scout separates the two jobs that a single model blurs together. A media engine does the grounded reading, the transcript, the on-screen text, the frames across the runtime, and it lays out a candidate timeline that is just the record of what is there. The Scout's governed model then does the judgment, choosing which moments matter and writing them up, with each choice anchored to a real timecode. If the engine cannot read an asset, that is reported, not smoothed over.

The result is the same shift the Scout makes everywhere, applied to the hardest input it handles: from a recap you take on faith to a recap that carries its own sources. The rest of this paper is how that works.

01 / The mission

It reads what people skim

A Forge Scout does not run once on a single file type. Like every Scout, it owns a standing mission, turning recorded media into something readable, and it takes in the whole range of media that work actually arrives in:

  • Video. The hardest case, and the one we use throughout this paper. On a video it does three things at once: it runs speech to text for a timed transcript, it reads the text that appears on screen, and it samples key frames across the entire runtime so nothing in the back half goes unwatched.
  • Audio. A recording with no picture still gets a timed transcript, so a talk or a call becomes something you can read and point into.
  • Documents. The text is pulled out and read, so a slide deck or a report joins the same brief as the video it sat next to.
  • Images. A still is read for the text in it and described for what it shows, the same way a single frame of video is.
  • 3D models. An object you would otherwise have to spin around by hand is taken in as part of the same source set.

The point of reading every kind of media the same way is that a real source is rarely just one file. A launch has a video, a deck, and a set of stills, and the Forge Scout reads all of them into one grounded picture rather than handing you a separate skim of each.

Of all of these, video is the hardest, because it hides its important moments in hours of runtime. So that is the input we use to show how the whole thing works through the rest of this paper. If the idea holds for a long video, it holds for the lighter inputs, which are smaller versions of the same work.

02 / The approach

An engine grounds it, a model chooses

The central choice is to split the work in two. A media engine does the grounded reading and lays out, deterministically, every moment that could matter, a candidate timeline built straight from the transcript and the frames. The Scout's governed model then reads that timeline and makes the judgment a person would: which moments actually matter, and how to write them up. The reading is grounded; the choosing is governed. Neither half does the other's job.

Figure 1. Recorded media goes in on the left: video, audio, documents, images, and 3D models. A Forge Scout reads each one for a timed transcript, on-screen text, and key frames across the runtime. A media engine builds a grounded candidate timeline, the Scout's governed model selects the moments that matter and writes the brief, and a sourced, readable brief comes out on the right.
Figure 1 · Media in, a sourced brief out. The engine does the grounded reading and lays out the candidate moments; the governed model picks the ones that matter and writes them up; the brief that comes out is tied back to the source.

Two things make this more than a pipeline. First, the candidate timeline is deterministic, so it is the same honest record of what is in the source every time, not a fresh guess on each run. Second, when the governed model cannot be reached, or could only answer poorly, the run does not stall and it does not pretend. It falls back to a plain selection straight from the engine, and it labels that fallback for exactly what it is, rather than dressing it up as a real choice. A recap should never quietly downgrade from judgment to a default and let you believe nothing changed.

That same split, an engine that grounds and a governed model that chooses, drives both of the Forge Scouts built on it so far. They are less two programs than one Scout pointed at different work: one at a pile of mixed sources about a single subject, the other at a single long recording. The next two sections take them in turn.

03 / Forge Analyst

It reads the whole data room, and writes one brief

The first of the two Forge Scouts is the analyst. Point it not at a single file but at a data room, the mixed pile of material that gathers around one subject: a slide deck, a recorded earnings call, a product demo, a set of stills, a 3D model. On their own, each is a separate thing to open and squint at. The analyst reads all of them and writes one brief.

Figure 2. A mixed data room on the left holds many assets about one subject: a slide deck, an earnings call, a product demo, product stills, and a 3D model. A Forge Analyst reads each asset on its own terms, getting the deck's text, a timed transcript of the call, a transcript and frames from the demo, and descriptions of the stills and the model. On the right it writes one sourced dossier where each line is tagged to the asset it came from, a closing line synthesises across the sources, and any asset it could not read is named as unread rather than filled in with a guess.
Figure 2 · Mixed sources in, one sourced dossier out. Each asset is read on its own terms; the brief that comes out carries every line back to the asset it came from, and names anything it could not read rather than guessing.

What makes the analyst an analyst, rather than a converter that turns five files into five summaries, is how it holds the whole set at once:

  • It reads each asset on its own terms. The deck gives up its text, the call becomes a timed transcript, the demo gives a transcript and the frames of what it showed, the stills and the model are read and described. Nothing is forced through the wrong reader.
  • It does not collapse into the biggest file. A long video in the pile does not drown out a one-page memo. Each asset is weighed for what it actually says, not for how many minutes or megabytes it runs to.
  • Every claim carries its source. The brief is grounded in the extracted content and tagged back to the asset it came from, each source named and listed, so a reader can go straight to the original and check it.
  • It opens by naming what it had. The brief starts by saying which assets yielded something usable and which did not, so a reader knows the ground the analysis stands on before reading a word of it.

And it ends where a careful analyst would, with a line drawn across the sources. Not a fifth restatement of the deck, but the thing only someone holding all of them at once can say: where the call and the deck agree, and where they do not. That closing synthesis is the part a stack of separate summaries can never give you.

04 / Forge Keynote

It watches the whole event, and writes the recap

The second Forge Scout is the keynote. Where the analyst reads a pile of assets about one subject, the keynote reads one long asset from end to end. Point it at a recorded launch event, a long one, and it returns a blog-style recap of what was announced. Not a transcript, and not a single paragraph that flattens a two-hour event into a sentence, but a structured page a reader can actually use.

Figure 3. A long recorded launch event goes in on the left. The Forge Keynote variant finds the reveals across the runtime and returns a blog-style recap on the right: one section per announcement, each with a plain-language summary grounded in what was said, a timecode pointing back into the source, and a short embedded clip of the moment it was shown.
Figure 3 · A keynote becomes a per-section recap. One section per announcement: a plain-language summary grounded in what was said, a timecode back into the source, and a short clip of the moment it was shown.

Each section of that recap carries four things, and the four together are the point:

  • One section per announcement. The Scout finds the reveals in the runtime and gives each its own section, so a reader skims the launch the way they would read a good write-up of it.
  • A plain-language summary. Each section is written in plain words, grounded in what was actually said at that point, not in what a launch like this usually says.
  • A timecode into the source. Each section points back to the moment it came from, so a reader can jump to that exact spot in the recording and check it. The page is built to be read that way: a contents rail down the side, sized by how much happens in each part, a still from every section, and a strip of frames from across the whole recording that you can click straight into.
  • A short embedded clip. Each section carries a short clip of the moment it describes, cut from the recording itself, so the proof is right there next to the words.

Read that list back as a sequence of actions and it is striking how much it is: it watched a long event, found the reveals, cut a clip for each one, wrote a grounded summary for each one, and handed the finished page to the gate. That is the work of an afternoon for a person, done end to end and handed over as one readable page.

Live · Forge Keynote

Apple WWDC 2026: Introducing Siri AI and more

Click to see the Forge Keynote summary.
05 / The discipline

It reads the heavy source once

One part of this is worth a paragraph on its own, because it is the difference between a thing that works in a demo and a thing that is kind to the machine it runs on. A long video is a large file. Cutting a dozen clips from it the obvious way would mean opening and reading that whole heavy file a dozen separate times, once per clip. That is wasteful, and on a long enough event it is the difference between a brief that arrives and one that does not.

So the Forge Scout does it the careful way. It takes the source in once, addresses it by its content, and cuts every clip from that single ingested copy. The heavy read happens a single time; everything after that reuses what is already in hand.

Figure 4. Two ways to cut a dozen clips from one long video. On the left, the naive way reads the whole heavy file once for every clip, which is wasteful. On the right, the Forge way ingests the source once, addresses it by its content, and cuts every clip from that single ingested copy.
Figure 4 · Ingest once, cut everything from one copy. The naive way re-reads the whole heavy file for each clip; the Forge way reads the source a single time and cuts every clip from that one copy.

This is not a flourish. It is the kind of care that decides whether a Scout can be pointed at a real two-hour event and still finish, or whether it quietly falls over on anything longer than a clip. Doing the heavy work once, and once only, is what lets the rest of the run be quick.

06 / The gate

It does not publish for itself

The Forge Scout produces the page. It does not get to decide that the page ships. That decision belongs to a gate, and the gate is the same one every Scout's output passes through, not a special exception made for media.

  • The Scout's job ends at a governed deliverable. It makes the recap, with its summaries, its timecodes, and its clips. Then it hands that off rather than pushing it live itself.
  • The page passes a publication wall. The same gate that checks every Scout's output checks this one. A page that is thin, that hides what it could not read, or that is not really finished does not get through. The check runs before anything ships, and if the gate itself cannot run, nothing ships rather than slipping past it.
  • A separate step does the publishing. Delivery is handled apart from the Scout that wrote the page, so producing and publishing are never the same hand.
  • It is confirmed genuinely live. Publishing reads the page back after it ships and checks the live page is the real one, not a placeholder or a half-built shell. If that read-back does not come back clean, the run does not call the page published. It would rather report that nothing shipped than leave yesterday's page up and tell you it is today's.

The reason for the split is the same reason a newsroom separates the writer from the editor who runs the piece. A Scout that could approve and publish its own work would be back to asking you to take it on faith. Keeping the gate separate is what makes the published page something you can trust without re-doing it.

07 / The honesty rails

It would rather say so than invent

Faithfulness is not something you add at the end. It is a set of refusals built in. A Forge Scout will not:

  • Use footage that is not the source. Every clip is cut from the actual recording by the engine. It does not reach for stock footage to stand in for a moment, because a clip is supposed to be the proof, not an illustration.
  • Invent content it could not read. When the engine cannot pull usable content out of an asset, the brief says so plainly. A gap is named as a gap, never filled with a plausible-sounding moment that was never there.
  • Pass a fallback off as a choice. When the governed model is unavailable and the run falls back to a plain selection, that is labelled as a fallback, so a reader is never misled into thinking real judgment was applied when it was not.
  • Pretend it has earned more trust than it has. Every run is gated, and a page only ships after it clears the wall, the same earn-it discipline every Scout follows. It is described that way on purpose.

We are deliberate about that earn-it posture. The Forge Scout runs and produces real, sourced pages, and when a run clears the gate it publishes one, but it is proving itself before it is turned loose, the same earn-it discipline every Scout follows. What it has shown is the part that matters most: that turning hours of recorded media into a sourced, readable brief, with every line pointing back to the moment it came from, is not a slide. It is something that can run on its own and be useful.

A lone AI tool
  • Skims a long video and writes a tidy summary.
  • Smooths over the parts it lost the thread on.
  • You cannot tell what came from the source.
  • Reaches for stock footage to illustrate.
  • To trust it, you watch the whole thing again.
A Forge Scout
  • Reads the transcript, the screen, and the frames.
  • Names a gap it could not read as a gap.
  • Every line points to a timecode in the source.
  • Cuts every clip from the recording itself.
  • The proof travels with the brief.
08 / What you can check

It hands you the means to check it

An honest brief is a good start. A brief you can check yourself is better. The Forge Scout ships the page with its proof attached, so trust is something you verify rather than something you extend.

  • Every clip carries a passport. Each clip is fingerprinted and signed against the exact moment in the source it was cut from. The published page lets you re-check that fingerprint yourself, in the browser, with one click. If the clip is the real source moment it confirms; if a single frame has been altered it goes red and names what changed. A clip whose origin cannot be checked is marked as unverified, in plain sight, never waved through as genuine.
  • The brief scores its own faithfulness. Before a page ships, a second pass goes line by line looking for any claim not tied to a real moment in the source. If it finds one, the page does not ship. What does ship carries a faithfulness score: how much of the brief is provably sourced, printed on the page rather than promised.
  • It flags footage it cannot vouch for. The incoming recording is checked for the signs of tampering, and any stretch the engine cannot stand behind is marked. It never calls footage authentic, because it cannot prove that. It tells you what it checked and what it could not, so a doctored source cannot quietly become a confident claim.
  • You can ask the recording. Beyond the written recap, you can put a question to the recording itself and it takes you to the exact moment that answers it, clip and timecode in hand. When the answer is not in the source, it says so, rather than reaching for one.

None of these ask for more trust. Each one hands you a way to withhold it until you have checked for yourself. That is the whole posture of the Forge Scout made literal: not a brief you believe, a brief you can audit.

Figure 5. A clip in the brief carries a passport: a fingerprint of the exact source moment it was cut from, signed and source-bound. You click verify and your own browser re-derives the fingerprint. It either matches, confirming this is the source at that second, verified not asserted, or a frame was altered, which is named and never shipped as real. The brief also carries a faithfulness score: the share of it that is provably sourced.
Figure 5 · The page proves itself. Every clip carries a fingerprint of the exact source moment; one click re-derives it in your browser and confirms it or flags a change. The brief also prints how much of itself is provably sourced.
09 / Across your library

It sees what no single recording can show

A brief reads one source well. The more useful thing is what only appears when you hold the whole library at once. As your recordings accumulate, the Forge Scout reads across them, and surfaces what no single brief can contain.

  • It catches a contradiction across two recordings. When one source says a thing and another, weeks later, says something that cannot also be true, it puts the two side by side, both clips, both timecodes, and names the tension. It does not pick a side or smooth the disagreement away. It shows you the conflict and leaves the judgement to you.
  • It traces how a decision was made. Across a run of meetings, it can lay out how a decision actually moved: where it was first raised, the moment that changed it, where it was settled, each step a clip you can replay. Where the record has a gap, it says so, rather than inventing the line that would join them.

Both of these need a library to read, and they get sharper as more of it is read. Both stay strictly inside your own walls: one client's recordings never cross into another's. The point is not that the Scout remembers more than you do. It is that it can hold all of it at once, and take you to the single place where two things do not line up.

Figure 6. Your library of recordings goes in on the left. The Forge Scout reads across them and surfaces two things a single brief cannot contain. A contradiction: one source at 47:10 says a feature is still in beta while another at 12:20 says it is generally available, both clips shown side by side, neither chosen. And a decision traced across meetings: raised in call A at 22:10, changed in call B at 08:30, settled in call C at 41:55, each step a clip, with any missing link named as a gap rather than invented.
Figure 6 · What only the whole library shows. Reading across your recordings, it surfaces a contradiction between two sources (both clips, both timecodes, neither chosen) and traces how a decision moved across meetings, naming a gap rather than inventing it.
10 / The bigger picture

A reader that gets sharper over time

None of this is built only for keynotes. A Forge Scout is one of a wider fleet of Scouts, so the same handful of qualities carry over to media:

  • It learns as it goes. It remembers which briefs proved useful and gets sharper each time, so the way it reads improves with every recording instead of staying fixed.
  • It works as a team, not a soloist. A Forge Scout is not a lone program but one member of a wider mesh of Scouts, sharing what it learns with the rest, so a useful find belongs to the whole mesh rather than to a single agent that can discover it once and lose it.
  • It shows its work. Every section comes with the timecode and the clip behind it, so the recap is something you can audit, not something you have to take on faith.
  • It uses what you already have. Beyond a single file, it can read the documents, stills, and recordings your team already works with, so the brief reflects the whole source rather than one slice of it.
  • Your information stays yours. One client's media never crosses into another's, and you decide what is ever shared outside.

And it improves within the limits you set, and stays inside them. It learns from briefs that were checked and held up, and still has to clear the same publication gate every time before a page ships, never by helping itself past it. That is the difference between a tool you maintain and a teammate that grows.

11 / Conclusion

A recap you can point at

Beyond the Agent argued that the Scout (persistent, collective, self-improving, accountable) is the unit that turns agentic AI from a pile of soloists into something that compounds. The Forge Scout is that argument pointed at the messiest input there is, where it is easiest to be fluent and hardest to be faithful.

It does not replace the watching. It does something more useful: it does the watching for you, completely, and hands the result over with the sources attached. An engine that reads the whole runtime. A governed model that chooses what matters and labels itself when it cannot. A gate that decides whether the page ships. A recap where every line points back to the moment it came from. That is the difference between an AI that summarises a video and one that can show you where every word came from.

A tool summarises a video.
A Forge Scout sources every line.

For anyone whose work means turning hours of recording into something a reader can trust, that difference is the whole game. The summary was never the hard part. Being able to point at where it came from was.