Vetrina

Italian · the shop window · for coding agents · MIT

Chat only ever grows.
A page is rewritten in place.

An hour with an agent leaves you a pile. In the order it happened, not the order it matters, and none of it addressable.

A vetrina page is scoped to one task and rewritten as the work moves. One screen, current state, shaped like the job: a migration gets a table, an incident gets a timeline.

teach your agents, once. any agent, not just Claude Code
open the window. run it anywhere, once, for every project

30 seconds · beat sheet and download


Ten pages · ten different shapes

Every one of these was built by an agent that had never seen the others.

Each was given the finding its page had to carry and never the layout, because a brief that describes markup has taken the agent's job. They share one stylesheet that owns the frame and the spacing and deliberately owns nothing in the middle: there is an axis, a track, a bar and a cell, and no component that draws a whole timeline.

They came out between 113 and 579 lines, with eight distinct combinations of drawing primitives and three that use none at all. Real projects, invented data: Gregario, Luxifare and Officina are mine and the findings about them are made up. The one with an outage on it uses a fictional company. These are live. Click one.

Running below:
real pages, not screenshots

dual-track timeline

The wrong theory lasted 19 of the 26 minutes

Two clocks on one axis. The time lost to the wrong cause is drawn as a hole, not described in a sentence.

Open the page→
contribution waterfall

Nxe5 wins the pawn and mortgages the king

Thirty-two named concepts scored the move. Four decided it, and two of those pull against each other.

Open the page→
annotated artifact

Ten invariants have a file. One has a habit.

The contract rendered as the document it is, each rule marked in place against the code that enforces it.

Open the page→
distribution

The median said ship it

Ten thousand voice turns. The middle is fine and the entire story is in the tail.

Open the page→
branching trace

Sync held for 47 writes, dropped the one that mattered

From signal loss to reconciliation, and the single fork where last-write-wins picked the wrong side.

Open the page→
proportional map

Nine lines reach every device

Forty-one config files sized by length and marked by reach, so the mismatch is a picture.

Open the page→
slice heatmap

+22 Elo overall, −24 in closed positions

A clear net gain hiding one class of position where it got worse. The aggregate cannot show this.

Open the page→
side by side

Three columns drifted, one on purpose

The same table declared twice, aligned line for line, so a difference is a break in alignment.

Open the page→
parking lot

Nobody ever came back to the backup question

Five asides dropped into chat while an agent worked, and what became of each. Two still unanswered.

Open the page→
no chart at all

One question, one screen

A hundred and thirteen lines. Most of what the skill says is when not to build a page, and this is what that looks like.

Open the page→

The problem is accumulation

The same hour of work, twice.

Left: the pile, in the order it happened. Right: the same hour, scoped to the task and rewritten in place. One of them gets longer every time the agent does something. The other does not.

Chatonly grows
I've finished the migration pass across the router files. In total I touched 41 files. Most of the changes were mechanical: useHistory became useNavigate, Switch became Routes, and the component prop became element. There were six places where I had to make a judgement call, mostly around nested routes and the custom PrivateRoute wrapper, and I'd suggest you look at those. Two of them use history.block, which has no direct v6 equivalent, so I picked an approach but you should confirm it. I could not finish three files. UnsavedChangesPrompt.tsx still imports the deleted Prompt component, which is why the build does not currently compile. On tests: 212 of 216 pass. The four failures are all in the auth redirect suite. Three of them look like the same referrer to from rename, and the fourth might be a real regression I introduced…
Vetrinarewritten in place
■Build does not compileBlocked
■UnsavedChangesPrompt.tsxBlocked
▲history.block has no v6 equivalentJudgement
▲PrivateRoute wrapper rewrittenJudgement
◆Redirect precedence, needs a decisionNeeds you
●32 mechanical conversionsDone
●212 / 216 tests passingDone

Sorted by how much of your attention each deserves. Filterable. The thirty-two boring ones collapse into a tally. Same information, but only one of them you can act on.


The hard part

Knowing when not to build a page.

Rendering structure turned out to be the easy half. Ten agents did it well with no kit and no house style. The half that decides whether any of this is worth having is the judgement call before the markup: does this reader have to hold more than about three things in their head at once, and if not, say it in two sentences and build nothing.

That judgement ships as a skill. It carries the test, ten named shapes to choose between, and the four constraints. Most of what it says is when to stay quiet.

  • A vetrina built for every session is a vetrina nobody opens. One unnecessary page costs you the odds anyone opens the next one.
  • Pick the shape before writing markup. A console bolted onto a problem that wanted a timeline has thrown away the whole premise.
  • Never the only place an answer lives. If they asked a question, answer it in chat too. The page is where the detail goes so the words can be short.

Given four real repositories and no mention of vetrina, it picked the right shape every time it built a page, and correctly built nothing for a question that wanted one line.

// when the reader would otherwise hold
// more than three things at once

mixed states, some needing you   triage queue
many items, secretly few causes  ranked clusters
options against hard limits      filterable list
a path, and where it breaks      branching trace
a number that moved, and why     chart + attribution
the artifact, defects in place   annotated artifact
paths a human must choose        decision matrix
a sequence where gaps matter     dual-track timeline
several agents at once           urgency-sorted lanes
asides dropped mid-task          parking lot

// and most of the time, none of them.
// the answer is a sentence. say it.

How it works

A directory, a convention, and a thousand lines of plain Node.

Anything that can write a file can publish to it. Claude Code, Codex, Cursor, Aider, a cron job, a CI run, a shell script. No SDK, no integration, no per-tool plugin.

Run it once

It binds your tailnet if you have one and loopback otherwise, refuses anything public, steps to the next port if yours is busy, and prints a QR code. No config, no account, no flags.

npx vetrina-cli

An agent writes a file

One self-contained HTML file into a space inside the window. That is the entire API. No registration, no manifest, no second file.

~/vetrina/<space>/page.html

Your tab updates

A watcher and server-sent events, injected into every page. The agent rewrites as work lands; the page in your hand changes without a refresh.

open on your phone, in another room

Evidence, not assertion

10/10
Ten pages · ten independent agents
Zero network requests · zero console errors
Judged against a rubric written first

The threshold was seven. The caveat matters more than the score.

Every agent was handed an already-structured brief. So what this proves is that agents render structure well when they have it, which answers the question about the design kit, and settles that the kit should be extracted rather than invented.

It does not prove that an agent notices when a page beats a paragraph. That needed its own test: four agents, four real repositories, no prompt mentioning vetrina at all. They picked the right shape every time they built a page, and correctly built nothing for a question that wanted one line, which is the harder half. One agent finished a twenty-finding audit and typed all of it into chat instead, and fixing that turned out to be a wording problem in the skill's own description rather than anything structural.


Where this goes

What exists today, and what is only an idea.

Everything above this line is built and shipped. Everything below is not, and is written down so it is obvious which is which.

Pages that answer back

A page carries a small block of structured data describing what it shows, including bounded choices it offers. You tap one from a phone and the agent resumes.

Blocked on a real question: a write endpoint would let anything that can reach the daemon act on your behalf, which is a far bigger claim than serving files.

Expiry and unread state

Pages age out on their own, the ones you have not seen are marked, and a page worth keeping gets promoted rather than quietly surviving.

Wanted once a window has enough in it to feel crowded. Not yet true.

Agents reading each other

If the data block exists, one agent can read what another published, and the window becomes where work coordinates rather than only where it is displayed.

The largest claim in the project and the least tested. Two agents and one handoff would prove it; fifty would not.

A relay, if the tailnet wall is real

Reading your agents from a phone with no network setup at all, across machines.

Deliberately last. Find out first whether loopback and a QR code already cover it for most people.