Italian · the shop window · for coding agents · MIT
An hour with an agent leaves you a pile. In the order it happened, not the order it matters, and none of it addressable.
A vetrina page is scoped to one task and rewritten as the work moves. One screen, current state, shaped like the job: a migration gets a table, an incident gets a timeline.
30 seconds · beat sheet and download
Ten pages · ten different shapes
Each was given the finding its page had to carry and never the layout, because a brief that describes markup has taken the agent's job. They share one stylesheet that owns the frame and the spacing and deliberately owns nothing in the middle: there is an axis, a track, a bar and a cell, and no component that draws a whole timeline.
They came out between 113 and 579 lines, with eight distinct combinations of drawing primitives and three that use none at all. Real projects, invented data: Gregario, Luxifare and Officina are mine and the findings about them are made up. The one with an outage on it uses a fictional company. These are live. Click one.
Running below:
real pages, not screenshots
Two clocks on one axis. The time lost to the wrong cause is drawn as a hole, not described in a sentence.
Thirty-two named concepts scored the move. Four decided it, and two of those pull against each other.
The contract rendered as the document it is, each rule marked in place against the code that enforces it.
Ten thousand voice turns. The middle is fine and the entire story is in the tail.
From signal loss to reconciliation, and the single fork where last-write-wins picked the wrong side.
Forty-one config files sized by length and marked by reach, so the mismatch is a picture.
A clear net gain hiding one class of position where it got worse. The aggregate cannot show this.
The same table declared twice, aligned line for line, so a difference is a break in alignment.
Five asides dropped into chat while an agent worked, and what became of each. Two still unanswered.
A hundred and thirteen lines. Most of what the skill says is when not to build a page, and this is what that looks like.
The problem is accumulation
Left: the pile, in the order it happened. Right: the same hour, scoped to the task and rewritten in place. One of them gets longer every time the agent does something. The other does not.
Sorted by how much of your attention each deserves. Filterable. The thirty-two boring ones collapse into a tally. Same information, but only one of them you can act on.
The hard part
Rendering structure turned out to be the easy half. Ten agents did it well with no kit and no house style. The half that decides whether any of this is worth having is the judgement call before the markup: does this reader have to hold more than about three things in their head at once, and if not, say it in two sentences and build nothing.
That judgement ships as a skill. It carries the test, ten named shapes to choose between, and the four constraints. Most of what it says is when to stay quiet.
Given four real repositories and no mention of vetrina, it picked the right shape every time it built a page, and correctly built nothing for a question that wanted one line.
// when the reader would otherwise hold
// more than three things at once
mixed states, some needing you triage queue
many items, secretly few causes ranked clusters
options against hard limits filterable list
a path, and where it breaks branching trace
a number that moved, and why chart + attribution
the artifact, defects in place annotated artifact
paths a human must choose decision matrix
a sequence where gaps matter dual-track timeline
several agents at once urgency-sorted lanes
asides dropped mid-task parking lot
// and most of the time, none of them.
// the answer is a sentence. say it. How it works
Anything that can write a file can publish to it. Claude Code, Codex, Cursor, Aider, a cron job, a CI run, a shell script. No SDK, no integration, no per-tool plugin.
It binds your tailnet if you have one and loopback otherwise, refuses anything public, steps to the next port if yours is busy, and prints a QR code. No config, no account, no flags.
npx vetrina-cli One self-contained HTML file into a space inside the window. That is the entire API. No registration, no manifest, no second file.
~/vetrina/<space>/page.html A watcher and server-sent events, injected into every page. The agent rewrites as work lands; the page in your hand changes without a refresh.
open on your phone, in another room Evidence, not assertion
Every agent was handed an already-structured brief. So what this proves is that agents render structure well when they have it, which answers the question about the design kit, and settles that the kit should be extracted rather than invented.
It does not prove that an agent notices when a page beats a paragraph. That needed its own test: four agents, four real repositories, no prompt mentioning vetrina at all. They picked the right shape every time they built a page, and correctly built nothing for a question that wanted one line, which is the harder half. One agent finished a twenty-finding audit and typed all of it into chat instead, and fixing that turned out to be a wording problem in the skill's own description rather than anything structural.
Where this goes
Everything above this line is built and shipped. Everything below is not, and is written down so it is obvious which is which.
A page carries a small block of structured data describing what it shows, including bounded choices it offers. You tap one from a phone and the agent resumes.
Blocked on a real question: a write endpoint would let anything that can reach the daemon act on your behalf, which is a far bigger claim than serving files.
Pages age out on their own, the ones you have not seen are marked, and a page worth keeping gets promoted rather than quietly surviving.
Wanted once a window has enough in it to feel crowded. Not yet true.
If the data block exists, one agent can read what another published, and the window becomes where work coordinates rather than only where it is displayed.
The largest claim in the project and the least tested. Two agents and one handoff would prove it; fifty would not.
Reading your agents from a phone with no network setup at all, across machines.
Deliberately last. Find out first whether loopback and a QR code already cover it for most people.