/Documentation

How LLM TimeMachine works

Lock a question. Ask it to as many AI models as you want. Come back in six months, ask the same question again, and see exactly what changed. This page walks you through it.

What is a capsule?

A capsule is one question, locked so it can never change, plus every answer any model has ever given to it.

That locking is the whole point. If the question stays exactly the same, then an answer from today and an answer from next year can be compared honestly — the difference comes from the model getting better or worse, not from you having reworded the question along the way.

A capsule showing its question marked as immutable and locked. Full size
Once a capsule has been run, its question is sealed — the interface shows it as read-only from then on.

The one thing to know before you start

A capsule stays editable until you run it for the first time. That first run seals the question, the files attached to it, and your checklist — for good. Nothing can unlock them afterwards.

You can still rename it, recategorise it, edit your notes and change who can see it. But if you want to try a different wording, make a new capsule: that way you keep both, and you can compare them.

Create your first capsule

On your dashboard, click New Capsule. Only a title and a question are required — but the optional fields are what turn a one-off test into something worth re-running in a year.

The New Capsule form, filled in with a title and a question. Full size
The creation form. Notice the note under the question: it will be locked on the first run.

Title

The name you will recognise it by. You can change it whenever you like.

Prompt

Locked on first run

The exact question every model will be asked. This is the part that gets frozen, so it is the part worth taking time over.

System prompt

Locked on first run

Optional, under "Advanced". A standing instruction sent ahead of your question every time — a role to play, a format to follow, a limit to respect. It freezes with the prompt, so it is part of the sealed question, not a setting you tune later.

Category

Helps people find your capsule if you make it public. Optional, and changeable later.

Visibility

Private capsules are yours alone. Public ones show up in Explore and can be read by anyone. You can switch either way at any time.

Attachments

Locked on first run

Up to 3 PDFs or images, 10 MB each, sent to every model along with your prompt — handy for "summarise this report" style questions.

Expected markers

Locked on first run

Words or phrases a good answer should contain. If an answer misses them, the run gets a small warning so you can spot it at a glance. You decide whether all of them are required, or just one.

Writing a question worth locking

The best questions are the ones models answer differently. If they all say the same thing, you learn nothing. Be strict about the shape of the answer you want — ending with something like Answer with a JSON list and nothing else. gives you results you can actually line up next to each other. And note down what a right answer must mention: those become your checklist.

Two things the capsule fills in for you

A sealed question can still say something different to each model. Drop {{VAR:MODEL_ID}} into your prompt (or your system prompt) and each model reads its own name there; {{VAR:RUN_DATE}} becomes the day the run happened. Buttons under both fields insert them for you. Every result keeps a copy of the exact wording it was given, so nothing about the comparison becomes guesswork later.

The name a model reads is its exact build, not just its family. Vendors reuse one name across releases — the April and the July DeepSeek V4 Flash answer to names that look nearly identical — so the model is told which dated build it is, and the result records it. Pick a latest entry in the model list and you get the build it points to today: handy, but that pointer moves, which is why every result keeps the name of the one that actually answered.

The three kinds of capsule

Most capsules are chat capsules: you ask, the model answers, you compare answers. An agentic capsule is a different exercise — instead of answering, each model actually goes off and builds the thing you described. In a Relay, the models take turns on one piece of work, each improving the version before. You pick which kind at creation, and it cannot be changed later.

Chat Agentic Relay
You give itA questionA job to doA piece of work to write and rewrite
The model gives backAn answer, or an imageWorking files — a page, an app, a gameA better version of the last one
How longSeconds to a few minutesHalf an hour or moreSeconds to a few minutes per step
What it costsOnly the modelThe model, plus the machine it works onOnly the model, a little more each step as the text grows
Who can use itEveryoneInvited accounts only, for nowInvited accounts only, for now

Agentic and Relay capsules are still invitation-only

If you do not see a Chat / Agentic / Relay choice when you create a capsule, that is why — they are open to a small group while we make sure they are safe and predictable. Everything else on this page works for everyone.

Run it against a model

You use your own OpenRouter account, which you connect once in Settings. That gives you hundreds of models — the big names and the open-source ones — and you pay them directly, at their price, with nothing added on top.

The run panel: filter by model type, pick a model, and add the run to the timeline. Full size
Pick a model and press Add run to timeline. The price per million words is shown before you commit.

The request goes straight from your browser to the model. Nothing passes through us, which is good for your privacy — and worth knowing for one practical reason: the run lives in your open tab. You will see the answer appear as it is written, and a model that genuinely needs six minutes to think will get them.

Four things you can adjust

Thinking

Some models can reason through a problem before answering. You choose how hard they think, from a light pass to an exhaustive one — deeper thinking usually means better answers, more time, and a bigger bill. You can also let the model think privately and only keep its final answer.

Web search

Lets the model look things up online before answering, which matters for anything recent. It costs a little extra, and the result card tells you whether the model really searched or just answered from memory.

Canvas

Ask for a working web page instead of text. You can then open what each model built and click through it — the fastest way to see the gap between two models on the same brief. Nothing runs until you press play.

Answer length

A ceiling on how long the answer may be. Leave it on Auto for normal questions; raise it for long documents or code. A cut-off answer almost always means this was set too low, and you can pick your own default in Settings.

If you close the tab by accident

Your run will not vanish. We warn you before you leave while one is still running, and if it does get interrupted you come back to a clear interrupted card with a Retry button — never a blank result pretending everything went fine.

Agentic capsules

Here you are not asking a question, you are handing out a job: build me a page that does this. Several models take it on at the same time, each on its own, and you get back what they actually built. This one does not live in your tab — close it, come back an hour later, the work carries on without you.

Two ways to give the job

From scratch

Just your brief. The model builds from nothing — a page, a small app, a game — and you can open the result and use it right there.

On an existing project

Point at a public GitHub project (up to 200 MB) and ask for a change: add a feature, fix a bug, make the tests pass. Each model works on its own fresh copy, and what you get back is the change itself, not a page. Nothing is pushed to your GitHub, unless you set up a Relay that pushes.

What happens once you start

01

It gets a workspace

Each model receives its own private machine, walled off from everything else.

02

It works on its own

It reads, writes files and runs commands, checking its own work, until it decides the job is done. It can search the web and read any page when it needs to; each search is a small charge on your key. Every model gets the same tools, so the only thing being compared is the model.

03

The result is kept

Every file it made is saved, along with a full log of what it did and how long it took.

04

The workspace is destroyed

Nothing is left running once the job is over.

You are never left without a brake

A model working on its own for half an hour can waste a lot of money if it goes wrong. Five limits stop that happening:

  • A spending limit per model, set before you start: $1, $3, or your own figure (minimum $0.50). Work stops when the limit is reached.
  • A hard stop after 2 hours, whatever happens.
  • If a model goes quiet for 20 minutes, the job is ended rather than left to burn through the clock. The card warns you long before that.
  • A Stop button, so you can end a job the moment you see it going nowhere.
  • One comparison at a time: as many models as you like on one capsule, but not two capsules at once.

What you get back

An agentic result card showing cost, time, and a live preview of the page the model built. Full size
The result card: cost and time across the top, then the page the model built, its files, and a log of everything it did.

Preview shows the finished thing, working. Code lets you read every file. Web, when the model used the web, lists what it searched, the pages it read (and those that failed), and which of them the result links to; LLM TimeMachine records this itself, so it is not the model's account. Reasoning replays what the model did, step by step. And the download button packages the lot into a zip, so you can carry on from where the model stopped.

On a GitHub project, the card shows Changes instead: every file the model touched, line by line, with what the model says it did and a button to download the change as a patch. Checks shows what LLM TimeMachine did around the model, without taking its word for it: it installed the project, ran its build and tests before the model started and again after (or the check command you gave), and made sure the change applies cleanly to the original project. Warnings appear above the change when the model deleted files, touched deployment settings, changed the tests instead of passing them, or broke a check that used to pass.

An External label means the result pulls something off the internet to work — a font, an icon set. It is a note, not a failure.

Relay capsules Invitation-only

In a Relay the models do not each answer on their own: they take turns on the same piece of work, like runners passing a baton. The first model answers your prompt. Every model after it gets your prompt and the latest version, improves it, and says what it changed. It suits work that gets better with several passes — an article, a piece of code, a translation.

A tree of five versions: v1, v2, then two versions built on v2 (v3a and v3b, greyed and marked wrong), then v4, starred as kept. Full size
Each box is a version. v3a and v3b both start from v2; v3b was marked wrong, so it is greyed out. The highlighted branch is where the next step goes.

Adding a step

Pick a model and press Improve. You can add a short instruction for that step alone — Add numbers, Make it shorter. The model receives your prompt, the version it starts from, and the notes the earlier models left along that branch. It never sees the other branches.

Your prompt is the only one you write: the wording that asks each model to improve the last version is ours, and every step shows in full what it was sent.

Branches

You can continue from any finished version, so the work can split: v3a and v3b are two versions built on v2. If you do not choose, the next step continues from the version you kept, or else from the newest version that looks sound. That branch is the one highlighted in the tree. A List button shows the same versions as ordinary result cards.

Reading a version

A version opened on its Changes tab: the lines removed from the version before in red, the lines added in green. Full size
Click a version to open it. Changes shows, line by line, what it changed from the version before.

Changes shows what really changed, in red and green. Result is the version itself; Read as article opens it full screen, laid out like an article, with a check of every source link when you are signed in. Web, on a step launched with web search, lists what the model searched and the sources it cited, or says it did not search. Prompt sent is exactly what the model received. What a model says it changed is labelled as its own claim: the red and green lines are what to trust.

On a GitHub project

A Relay can also work on code. Pick Relay, then GitHub repo, and give a public repository. Each step runs like an agentic job: the model gets its own machine with the project as the steps before it left it. Their changes are replayed as commits, one per step, signed by the model that made it, so it can read who did what with git log. You give each step a role (add a feature, fix, review), and several models launched together make sibling versions you can compare.

A version shows the change to the code, the build and tests LLM TimeMachine ran before and after the model, what it searched and read on the web, and the brief it was sent. A step that breaks tests which passed before it is marked as broken and is not used as the next starting point unless you choose it. Warnings also say when a step removed lines the step before it had added, or deleted a file an earlier step created. Every change can be downloaded as a patch.

Pushing to GitHub

Tick Push each step to GitHub when you create the Relay, and every finished step becomes a commit on a relay/… branch of your repository, authored by the model that made it, for example DeepSeek V4 Flash (via LLM TimeMachine). It is automatic: there is nothing to click after each step. The choice is made once, at creation.

  • The first line of versions goes to relay/<capsule>/main, and each step built on the end of it moves that branch forward.
  • A sibling version, or a step continued from a version in the middle, gets a branch of its own, such as relay/6c21ce1b/step2-9f3e1a.
  • Branches only move forward: a version you delete or mark wrong keeps its commit on GitHub, and the Relay shows its real state.

The pull request stays your decision. On the version you want, Open a pull request opens one into the branch the Relay started from. It comes from a branch that stays at that version, even if the Relay moves on, and its description lists each model, its role, its change and its checks. You review it and merge it on GitHub, like any other pull request. If your repository builds previews (Vercel does, for every branch), the version also shows a Preview link: the site as that step left it.

What protects your repository

  • The models never get a GitHub token. LLM TimeMachine writes the commits itself, after the step is over.
  • It only ever writes branches whose name starts with relay/. Your own branches are never touched, and nothing is ever force-pushed.
  • A step that changes .github/ (your CI workflows) is not pushed, and the steps built on it cannot be either. Its change can still be downloaded.
  • GitHub must end up with exactly the code the model left in its machine. If not, the branch does not move.

If a push fails (GitHub unreachable, no access to the repository), the version says why and offers Push again; the steps before it that are not on GitHub yet go first.

Before you turn it on

Every pushed step starts your CI and your preview builds, with code written by models. Run the Relay on a fork or a separate repository, or tell your tools to skip relay/* branches (on Vercel, the Ignored Build Step; in GitHub Actions, branches-ignore). During the beta, pushes use the GitHub access of LLM TimeMachine's owner, so they only work on repositories it was given.

Steering it

Continue from here

Makes this version the starting point of the next step, even if it is not the latest. That is how a new branch starts.

Keep this version

Your pick. The next step starts from it until you choose another, and it is starred in the tree.

Mark as wrong

For a version that went off course. It is greyed out with everything built on it, and never picked as a starting point again. The mark can be removed.

Delete

Removes the version and every step built on it. You are told how many go with it before anything is deleted.

A version that came back cut off, empty, far shorter than the one before, stuck repeating itself or garbled is marked broken. It is never picked as the next starting point on its own; you can still read it, and still choose it.

Comparing fairly

Each step reads a longer text than the one before, so it costs more and takes longer — which says nothing about the model. So in a Relay a version is only ranked against the versions built on the same one, and Compare all puts the latest version of each branch side by side.

For now

Image models are not offered in a Relay. On text, web search is off unless you turn it on for a step, and a step runs in your open tab, like any chat run; a step on a GitHub project runs on its own machine and carries on if you close the tab. Only the capsule's owner adds steps: on a public Relay, other signed-in people will be able to add theirs later.

Understand your results

Every run is kept with far more than its answer. Months later, that is what lets you say why a model improved — it got cheaper, it got faster, or the provider quietly swapped it for a newer version.

A result card showing speed, size, cost and the model's answer. Full size
One result: speed, length, cost and the answer itself, with tabs for the finer detail.

What gets recorded

  • What it cost you, down to the fraction of a cent
  • How long it took to start answering, and how fast it wrote
  • How much text went in and came out
  • Which exact version of the model answered — providers update them silently
  • Whether the answer finished properly or was cut off

When a result is flagged

A model answering is not the same as a model answering well. We mark a run as worth a second look when:

  • It was cut off before finishing
  • It was blocked by the provider’s safety filter
  • The model declined to answer
  • It got stuck repeating itself
  • It missed the checklist you set on the capsule

What you can do with a result

If an answer was cut off, you can ask the model to carry on from where it stopped, or run it again with more room — which gives you a fresh result and leaves the original untouched. Results you would rather not count can be archived: they leave your timeline and your averages, but they are never deleted behind your back.

Compare models

Tick Compare on two or more results and open the Compare Studio. Speed, price and length line up as bars, so the trade-off is obvious at a glance.

Two models compared: one is four times faster, the other five times cheaper. Full size
A fast, pricey model against a slow, cheap one. Same question, very different bargain.

When you compare exactly two answers, you can highlight what changed between them, word by word — the quickest way to see whether a new model really said something different or just reshuffled the same points.

Two haiku answers side by side, with the differing words highlighted in each. Full size
Only the words unique to each answer are highlighted. Here, two models differ by three words.

To judge the results by eye, open Bento in the studio's header: every page, image or answer becomes a tile in a full-screen mosaic, with the model's name only on hover — or hidden altogether in blind mode. Click a tile to see it full size, then use the arrow keys to go from one model to the next. Only one page runs at a time, so even dozens of heavy results won't slow your browser down.

The same model, over time

Compare History pulls together every run of one model on this capsule, so you can watch it drift across months — which is the reason this whole thing exists.

Taking it away

Export the comparison as a web page to read or send to someone, or as a text file if you want to hand the whole thing to another AI for a second opinion.

One thing we deliberately do not do: crown a winner. Numbers alone cannot tell a good answer from a fast refusal — a model that replies "I can't help with that" in one second would win on speed every time. You get the figures and the answers side by side; the judgement stays yours.

Keys and privacy

Your OpenRouter key

Your questions go from your browser straight to the model — they never travel through our servers, and we never see them. Your key is stored in your browser and saved to your account so it follows you between devices. It is scrambled rather than protected by a password only you know, so treat it the way you would any key saved in a website: give it a spending limit on OpenRouter, and replace it if you ever stop trusting the computer you saved it on.

Never put a password or key inside a question

A question is sent to every model and locked forever — and if the capsule is public, anyone can read it. When your question needs a key, save it in Settings and drop in a placeholder instead, like {{SECRET:MY_KEY}}. Only the placeholder is stored; the real key is filled in at the last moment, and only for the websites you allowed it to reach.

Paste something that looks like a real key and we will spot it and offer to swap it for a placeholder in one click, before it can be saved.

The Capsule Secrets panel in Settings, where a key is saved with the websites it may be sent to. Full size
Settings is where keys live. Each one lists the websites it is allowed to reach — and nowhere else.

What "public" really means

Making a capsule public shows everyone its title, its question, your notes and all its results — answers, costs and speeds included. Search engines can find it too. Files attached to a private capsule stay private. You can go back to private at any moment, even after the capsule is locked.

Want to see one first?

Have a look at what other people have locked away. No account needed.