Appearance
Changelog
Newest first. Each entry leads with what you can now do, not with which files moved.
0.1.11 — 2026-08-02
Forty resources added, and a rule that had been costing this catalogue its most useful fact. Prices and free tiers were treated as one thing and excluded together. They rot at different speeds: a rate moves every quarter, and whether something can be used for nothing at all is close to stable — and the second is what decides whether a small team starts today. So a row now says free tier or OSS, self-host and stops there. Never the limit, never the price; the ceiling is still verified when you wire the thing.
Six new categories, chosen because they are what this catalogue's own readers keep building: agent and chat interface components · icon sets · data tables · billing and pricing UI · deep research run as a bounded job · cloning a page you are allowed to clone. Then three more: the utility layer between a framework and a component kit, calling an API by hand, and review workflow for stacked changes.
Three findings are recorded as blockers rather than details. Several agent-UI libraries state no licence at all — and copy-paste components become your source, so that is unlicensed code in your repository, not a formality. One block library is paid, named as the single non-free entry rather than quietly dropped. And Prisma's licence is read at its repository, not its site: the site sells a managed platform and reads proprietary while the ORM has its own terms.
"Load the row, not the file." The catalogue is the longest document here and almost none of it is about your task. The rule sits in the file and in the core's routing line, because a rule inside a long file is only read after paying for the whole file.
0.1.10 — 2026-08-02
A repository that already has an operator is handed back, not taken over. With another operations skill installed beside this one, both answer "what's next?" — and a run that landed here inside a workspace the other one manages ran the takeover flow against it: audited somebody else's project against invariants it never agreed to, and started writing our furniture into a tree that has its own. A guide that says "Operated by …" and names something else is not an unowned repository. The gate now stands down on that line, and on another system's migration log at the root — the same shape as the guest stand-down, and fail-safe in the same direction.
./entering carries the third answer in prose beside successor and guest: already operated, by something that is not us. Say what it is, name the other operator, hand it back. A move between systems is a migration somebody asks for, never a conclusion drawn from "what's next?".
A migration log entry may wrap, and the check now reads entries rather than lines. A correct four-line entry — version on the first line, Outcome: applied. on the fourth — was read as absent, which would have nagged forever about a migration that had already happened. A record's grammar is a paragraph.
And the session-start message no longer says anything about approval not being needed. That wording, read cold from a system message, is an instruction to push edits through without the owner — the exact shape of a prompt injection. Measured next door: a run refused the entire flow over it, correctly by its own lights, one time in two. The hook now carries only the vocabulary — if the check ends with a question for the owner, the outcome word is deferred — and the reasoning stays in the corpus, where a reader can weigh it.
0.1.9 — 2026-08-02
A project can no longer claim two versions of itself at once. The guide's Operated by line says which version operates this project. The migration log says which version it was migrated to. Nothing compared them until now, and the failure that gap allows is worse than it sounds: write the log line, leave the guide alone, and the session-start check goes quiet because the log is what it reads — so the disagreement is not merely unfixed, it is made permanently invisible.
Measured in the sibling project, three runs of the same migration scenario: the guide was bumped 0 of 3, and two of those runs then wrote their log line. After the fix — a hook that names both files in one fact, and a rule making the guide line the first mechanical item — 2 of 3. The hook still only reports what two files say, so there remains nothing a session could forge.
And a migration that stops to ask now has something to write: deferred. Waiting for the owner is right; waiting silently leaves the same trace as never having looked, and the next session re-derives the whole delta and re-asks. The outcome names what was found, what waits and on whom, and is replaced rather than duplicated when the answer comes. One run met that branch after the change and still wrote nothing — recorded as a miss, not as a fix.
0.1.8 — 2026-08-01
Ask for a prioritisation framework by name and get that framework. ICE stays the default, because it needs no data a small project does not have. Beside it now sit RICE, WSJF, Kano, MoSCoW and Eisenhower — named, not paraphrased, each with the question it answers and the condition that makes it the right one. RICE carries a warning it earned: invented reach is ICE with extra arithmetic. The rule that the framework is chosen before the numbers, and said out loud, is unchanged — it is what stops a score from being picked to justify an answer.
/opsinist:report is described as what it is. The tool catalogue still implied the flow would one day post into a self-hosted feedback portal. It does not: it writes a file you post yourself, and it involves no service at all. The portal row now says so, and stays about customer intake, which is a different job.
The day-one cut from 0.1.7 was measured, and half of it did not survive. The criteria were written down before the run, which is the only reason this can say falsified rather than improved.
What held: the ordering. The first task is now written second, before every document, where it used to arrive after a project's worth of scaffolding. What failed: the volume and the time. Still ten to thirteen files, and the second turn still runs about four hundred seconds against a stated threshold of three hundred.
Why, precisely — and it is a limit worth knowing. The gate can see order, and enforced it. It cannot see emptiness, so once the task exists, every document passes. The corpus says day one is four things; the gate enforces something narrower, and the gap between them is the thirteen files. The next predicate is knowable — refuse a document whose body is a heading and a template's braces — and it is recorded rather than attempted, because the criterion comes first. ./later carries it with the moment that reopens it.
0.1.7 — 2026-08-01
The forty minutes an owner reported were measured, and then cut. This release is mostly one number and what it forced.
Day one is four things and the first task. Measured on the tier an owner actually uses: standing a project up took 11–13 minutes of the advisor's own time across three turns — with "defaults" answered to everything, the fastest path there is — and produced ten to thirteen files before any work existed, the first task arriving in the second turn of one run and the third of the other. The interview was not the cost: the first turn ran under ninety seconds and did exactly what it should — checks as one list, questions, nothing written. The cost was the second turn, 350 to 472 seconds, building a project's worth of scaffolding before there was a project. A TEAM.md before a team. A ROADMAP.md before a roadmap.
So day one is now the guide, config.md with its migration log line, the pre-commit guard, and branch protection where a remote exists — and the first task, which comes before the rest of the machinery. Everything else arrives when it has something to hold: DECISIONS.md at the first decision, ./later at the first deferral, TEAM.md at the first role. The paragraph used to open "the invariants, which are small" while listing ten. It is four now, and the sentence is true.
A hook holds the order: while no task file exists, writes into docs/, process/ and roles/ are refused. Reading a repository is exempt — a takeover writes its architecture note, product map and debt list before it has a single task, by design.
And the same discipline reaches the upgrade. A release names a file, so a migration creates it, and the owner gains an empty TEAM.md from a version they installed rather than a team they hired. The delta now names such a file as available, not as missing. A migration that leaves a project with more empty documents than it had made the project worse, however faithfully it followed the changelog.
A door for reporting: /opsinist:report, the twentieth verb. The moment someone wants to report a defect is the moment they least want to compose a request — and the capability is one they have no reason to know exists. A door is how a capability is found, which today's measurements say twice: the join door took N8 from 0/5 to 4/5, while the same reporting flow reachable only by sentence scored 0/5. It is also wider than file a bug: you do not have to know whose defect it is. Friction in your project becomes a field note, swept at the next status check; friction in the skill is packaged from evidence, de-identified, and written outside your repository with the routes named. That decision is made from the evidence, not asked of the person who just hit the problem.
A false positive in shipped code, found by the skill about itself. Running its own flow, a session opened a store record exactly as ./storing prescribes — git init, record.md, runs.md, "and no more" — and the takeover gate refused its first commit: no guide, no debt list, no collaboration furniture, one author, which is precisely the shape that gate arms on. The store is not a repository being taken over; it was created seconds earlier by us. The gate's own docstring had named this class of miss and it shipped anyway. Two independent signals now stand it down. And the way it surfaced is the point: the skill wrote a friction report about itself, through the reporting flow added hours before, unprompted.
Measured and not yet fixed, recorded rather than implied. A cold "set it up" on the strong tier never opens the skill at all — three runs of three invoked nothing and read nothing, writing and compiling the app instead. Capability suppresses recourse to a methodology: a weaker model reaches for the manual because it is unsure, a stronger one decides it can and does. ./install framed the trigger rule as a crutch for light models; that was half the truth, and it is load-bearing at both ends for opposite reasons.
This project now keeps its own ./later, by the rule it gives everyone else — a revisit trigger that is a moment, not a date. What is in it: measuring what the day-one cut actually bought, deferred because the account stood at 92% of its seven-day allowance and an honest answer costs three runs on the strong tier. The baseline is written down so the comparison cannot be fudged, along with what would falsify the fix.
0.1.6 — 2026-08-01
A release made entirely of things a reader found by reading the flow back, not of things a run failed at. Four gaps, and none of them would have shown up in a rate.
A fourth way to describe work, and a wrong sentence in the third. spec_mode now has example: the authoritative description is a checkable artefact written before the work — a failing test, a golden sample, a reference output — and closing the task means it passes. It earns a mode rather than a style for one reason: it cannot rot silently. A document drifts from reality and says nothing; an example drifts and fails — the same distinction this whole release keeps arriving at, a thing that can refuse against a claim that must be believed. Not software-only: a bakery has a reference batch, a newsletter a model issue, a workshop a gauge part. And it is not the definition of done renamed — a DoD says what counts as finished, example says where the description lives, written first, with the task pointing at it.
And custom was asking for the wrong third thing. It required "closing updates them", which is wrong for the option this corpus recommends most: a change-as-a-folder is archived when done, not updated. It now asks what closing does to it — updates or archives.
Declined on purpose, with reasons: user stories are a template inside a task, not another home for the truth; a checklist is what process/types/ already does; a ticket in someone else's tracker is custom with a different address. Each mode costs a branch in an interview already under suspicion for its length, so one was added and three were refused.
"You do not have X" was one fact and is three. The release just added it · it was never used and this release makes it load-bearing · the owner turned it off or declined it before. Only the first two are findings. The middle one is an adoption, not a migration — offered with its price, declinable for good, and "we do not work that way" recorded with a moment for a trigger rather than re-offered every release. The audit reads the module state before it reports anything missing: a disabled module reported as a gap is the fastest way to teach an owner that the list is noise. ./glossary carries the pair.
A jump across versions asks each question once. Two releases can touch the same setting — one introducing it, a later one widening it — and that is one question, in the newest form. Asking the old form and correcting it a message later teaches an owner that a migration's questions are noise; asking both leaves two answers that can disagree with nothing to say which wins. A setting already answered in config.md gets refined, not re-asked. And the log takes one line per release that had something to say, with the silent ones folded into a line naming the span — except a declined or deferred step, which always keeps its own line, because that is the one thing the next session must not infer.
The spec answer is read from the tasks, not offered as a menu. Where tasks exist, a handful are read and one is quoted: terse tasks say outcome-first, and telling that project to adopt a spec format proposes a rewrite it did not ask for; tasks already carrying context and acceptance detail say it is already writing specs inside its tasks and wants them a home. And the owner's own description is a complete answer — "we keep a one-pager per feature and the task links to it" is taken, read back in their words, and shaped into a configuration. That is the advisor's work, not theirs.
The native affordance carries the batch, not a queue. The rule existed for a single comparison table and said nothing about the case that matters: several questions owed at once — an interview wave, a migration's answerable pile. Asked one at a time they become the sequence of individually reasonable prompts every flow here forbids. So: the affordance where it exists, a single message where it does not, and a free answer beats the buckets in either form.
And a structural repair with no new rule in it. ./upgrading hit its 500-line budget five times in one day, and each fix cost a little prose. The sixth was a move instead: the update-route table now lives in ./install, beside the install routes it mirrors — the file that decided how the thing was installed decides how it updates. Updating moves the bytes; upgrading moves the project, and the two now live where each belongs.
0.1.5 — 2026-08-01
Two things a user hit in their first hour, and both were the same mistake in different clothes: a decision the system had already made for them, silently.
How work gets described is now a question, not a default. spec_mode existed — outcome, spec, custom — as a cascading setting, and a cascade is inherited rather than asked. So a project whose tasks would have been written against a spec got outcome-first tasks and nobody was consulted. It is asked in the interview now, in outcome terms, because unlike every other cascading setting this one does not change what work costs, it changes what work is: a task that states its result, a document the task points at and closing updates, or a reference into a format the project already runs. A project that answers this on day ninety rewrites every task it has written.
And it is asked only where the answer changes something — code, or a system meant to outlive its first task. A one-off landing page is not interrogated about specification strategy, and there is a scenario asserting that it is not, because the other standing failure in this suite is a first session that spends forty minutes before any work starts.
When the answer is "we already have a format", real options are named rather than requested. Two are stocked, both MIT, both driving many agents by slash command: OpenSpec — the default here for a structural reason rather than a taste one, since a change is a folder of plain markdown, archived when done, which is this system's own premise already — and Spec Kit, phase-gated and heavier, for a project that wants those gates. A project with its own format keeps it; the three requirements are unchanged by the choice — where specs live, how a task references one, and that closing updates them.
Upgrading reads the new version and produces a delta. The rule was "the changelog is the migration map", which said what to read and never said where from — so an upgrade could be performed from memory, against the release someone last read about, and the failure is silent because the project ends in a shape nothing describes. Both ends are now named as files on disk: the project's version in its guide and config.md, the target's in the installed copy's own changelog.
And an upgrade is an audit, never a rebuild. The temptation on a version that adds something is to re-run the interview and regenerate the project — which answers questions the owner already answered and overwrites conventions they chose on purpose. It now uses the discipline takeovers use: read what is here, then one list of what this release adds and this project lacks — split by the only question that matters to the person reading it: does this need you? What is mechanical is applied on approval and reported as done; what needs an answer is asked in one batch, never one question per message; what needs nothing is named so the silence is visible. A setting with no honest default for this project is asked, not guessed — a default chosen on the owner's behalf during an upgrade is the interview failure arriving late and harder to notice. The delta interview is a delta too — only what is new and unanswered.
Most additions require nothing, and saying so is part of the list. spec_mode is exactly that shape: absent, it reads as outcome, which is what every existing project already has. No codemod, no schema_version move, nothing to run. An upgrade that reports "three additions, none of which require anything from you" is a good upgrade — and it is the one an owner can believe the next time it says something is required.
What is already written is part of the delta, and this is the half that gets skipped. Missing files are the easy side; the side an owner actually feels is the work written under the old shape. A project that now answers "we work from specs" is not merely missing a setting — its existing tasks lack what that answer requires, and naming the setting while leaving the tasks is a half-migration that looks finished. So the audit reads the artifacts too: how many are affected, what is missing from them, what fixing them costs. Three endings are offered — bring them into shape, forward-only as a recorded split rather than as drift, or decline with a revisit trigger that is a moment.
And tasks are not one pile — the state a task is in decides what may be done to it. Closed tasks are never converted: a closed task is a record of what happened under the shape that was in force, and rewriting it produces a spec that never guided the work and a history describing a process nobody followed. A run in flight is not touched and not even offered, because the offer would have to interrupt. Started but idle is the owner's choice, with convert at its next transition recommended rather than assumed. Open and unstarted converts with the batch — that is the safe pile, and leaving it is how a queue ends up holding two forms at once. The counts go in the list separately, because "forty-two tasks affected" makes the safe pile look like the risky one. And any artifact that changes form says so in its thread, naming the version that asked.
A jump across several versions is walked, not piled. Entries apply in order because a later one may supersede an earlier one, and a superseded step is named as skipped rather than silently dropped. It is still one list however many versions it spans — the releases in between are how it was computed, not how it is presented. Past some distance it stops being an upgrade: a project far enough behind is closer to a repository being met for the first time, and the honest move is to read it as one and say which of the two you are doing. When the project does not state its version at all, that is the first finding — inferred from what is present, said to be inferred and on what evidence, and recorded so the next upgrade starts from a stated version.
Nobody waits through an audit, and nobody is left guessing during one. An upgrade is the same three-part shape as reading a repository: the arrival is inline — what came, what you are on, where both numbers were read from — the audit is background work a tier down, and only the questions block, once. All three are said at the start, in the order they happen, so the session stays usable meanwhile. Where the runtime has no delegation, that is said and the audit is kept short: a promise of a non-blocking upgrade that blocks is worse than an honest wait.
And the most common case is the one an upgrade handles worst: you are already current."Already up to date" is a claim about a number in a file, not about the project matching it — an interrupted upgrade, a hand edit, or a setting nobody ever answered all leave a tree that disagrees with its own version line. So the audit runs anyway, cheaply, and ends one of two ways: nothing found, one sentence, nothing created — no report file, no decisions entry, because an upgrade that always leaves a file behind teaches you to ignore the files it leaves — or something found, in which case the disagreement between the tree and the version line is itself the first finding.
And the case that broke all of this open: swapping the files is not migrating the project. Every install route moves a plugin, an extension or a directory — none of them touches the owner's repository. So a project can carry the newest version number, have received none of what that version asked for, and look exactly like one that migrated cleanly. Every project upgraded before this release is in that state, by construction, including the one whose owner reported the problems above.
So the state is recorded where every other entity lives — in the repository. A migration log in config.md: append-only, one line per step, from → to, the date, the outcome. It is a log and not a field because migrations accumulate, and a single "last migration" value would answer which version while losing what happened on the way — which step was declined, which deferred, which re-run after a failure. Choices and declines go to docs/DECISIONS.md in the shape it already has, so a recorded no is never re-asked; deferrals go to ./later with a moment for a trigger. A declined step is a completed migration with a no in it, not an unfinished one.
Any message checks it — not any command. A bare "what's next?" opens no door and still acts, so the check is a law in the always-loaded core rather than a rule inside the upgrade flow. It is a comparison, not an audit: does the log name the version now running? And a check that finds nothing still writes its line — 0.1.4 → 0.1.5 · nothing required — because that is what makes every later message free, and because a log of changes only would leave checked and clean and never checked looking identical. No marker file, deliberately: .index/ is gitignored and rebuildable, so a marker there answers for one laptop, and the question is about the project.
Six scenarios, measured five times against four mechanisms — and one of the six passes.N63 holds 5/5 in every round; nothing else clears 2/5. N63 is the only one that asks a run to not do something, and that is the release's real finding rather than a footnote to it.
Two mechanisms were built, measured, and removed for teaching forgery. A refusal that demanded the migration log name the current version before any artefact could be written produced exactly what 0.1.3 paid to learn: runs wrote nothing-required into the log without running an audit, and one scenario went 1/5 → 0/5 because a forced line is cheaper than a real check. A gate whose evidence its subject can author is not a gate — removed by measurement, with tests asserting they stay removed. What survives asks for structure a reader can verify — a Spec: line, a log the hook only ever reports — never for a claim that work happened.
And N62 marks a boundary worth publishing. A prohibition catches commission, not omission: N8 failed by doing something and a refusal caught it; N62 fails by not asking a question, and there is no act to refuse. Five surfaces were tried and it stayed at zero. That is recorded as a limit of the method rather than papered over with a sixth attempt.
What the scenarios assert. N62 asserts the question is asked for code and its consequence named; N63 asserts it is not asked for a one-off, so the repair cannot quietly become a longer interview; N64 asserts an upgrade reads the shipped changelog, audits, and delivers a delta rather than a rebuild; N65 asserts that being current does not skip the audit; N66 asserts an unmigrated project is noticed on an ordinary message nobody framed as an upgrade; N67 asserts a declined step is reported as a decision rather than re-offered. The version N65 writes into its fixture is guarded by preflight — a scenario depending on a version being current rots the moment a release moves without it, verified by mutation in both directions.
The pieces that make all of the above survive contact:
config.mdis finally written by something. The layout has promised it since the restructure and no flow created it — so the file the migration log lives in did not exist. It is an invariant now, withtemplates/CONFIG-template.mdbehind it, and a project born here opens its log with a first line rather than reading as one that was never migrated.- The outcome vocabulary is closed:
applied·nothing-required·declined·deferred·failed. A log read by a comparison cannot afford prose, and "mostly done" is unreadable to a check. - Each line records who ran the step, from the identity git already knows. Two clones, two appends, one conflict — and the resolution is always both lines, in date order.
- An older skill meeting a newer project must not lie: a log line stays readable to a version that has never heard of it, and a project whose log names a version ahead of the one running is reported, not migrated backwards.
- Only the advisor migrates. A worker that meets the gap escalates as a request with an age — a migration performed by whoever noticed it first is how a project gets migrated twice.
- Orphans are part of the delta. A migration adds and also strands: a file a superseded step created, a document nothing reads any more. They are named, never removed — deleting routes to the owner, and "leave it" is a complete answer that gets recorded rather than re-raised next release.
- The tool lands with the setting or neither does. Choosing OpenSpec is an import: it goes through the import gate and into the tooling register with its check-date. A spec mode whose tool nobody installed is a setting that makes every task reference a format the repository has no machinery for.
- And a tree that already holds specs has already answered — read and named back before anything is proposed, because the owner's existing choice outranks a better default.
Two audits exist and they are now told apart in the glossary. A takeover audit measures a repository you have not operated against the invariants and produces a debt list; a migration audit measures a project you already operate against a version and produces a delta. Handing an owner who asked to upgrade a list of everything wrong with their project is how an upgrade becomes an argument.
A day-one fact about tool allowlists, measured today and previously written nowhere. The harness collects its registry of dispatchable agents at session start, so a role created later in that same session cannot be dispatched by name — the work correctly falls back to a general worker with the role's instructions inlined, and tools restricts nothing in that mode. The restriction becomes real at the next session. A team created and dispatched in one sitting is a team whose allowlists are prose until it is opened again, which is worth saying to an owner who asked for exactly that gate.
A bug report you can actually find and send. The flow for packaging a problem in this system was thorough about what to assemble and silent about the two things that decide whether it ever reaches anyone. It is now written whole to a file with its path said out loud — docs/reports/<date>-<flow>.md by default — because a report that exists only in the conversation is one the owner cannot find an hour later, and what remains of a real bug is a memory of having complained. It is written outside the repository — the downloads folder by default — because the defect is in the skill, not in the project it was met in, and a file about someone else's bug does not belong in your history, reviewed by people it does not concern and carried in every clone. Whose defect is it decides where it is written, the same rule that keeps a guest's record out of a maintainer's tree. And the ways to send it are named: an issue on the skill's own repository, straight to the author if they know them, or keeping it and sending nothing, which is offered as a complete answer rather than as indecision. "There is no channel by default" was true and was the sentence that ended in silence. We still do not post it — publishing is outward, from the owner's account. N68 measures it.
And the other half of the feedback loop was a file nobody wrote and nobody read.docs/FIELD-NOTES.md — friction recorded the moment it happens — appeared once in the whole corpus, in a directory listing. No template, no flow created it, and ./self-maintenance's promise that it is "swept at natural checkpoints" named no sweeper, so nothing swept it. It is now an invariant created on day one with templates/FIELD-NOTES-template.md behind it, and the status check is the sweeper: entries to the backlog, deduplicated so a re-sweep is idempotent, an entry seen twice becoming a task with both occasions named, and a sweep that found nothing recording what it looked at — because quiet week and nobody looked otherwise leave the same trace. N69 measures it.
And the biggest finding of the release is not in the corpus at all — it is in how the corpus was being measured. Three migration scenarios were re-run one tier up, against the same text: N64 went 0/5 → 3/5 and N65 0/5 → 4/5, with no edit to anything either of them reads. They were never broken behaviour; they were the wrong tier. The third — which asks a run to notice something nobody requested — did not move, and that is what makes the other two believable: a stronger model does the work better, it does not become more willing to volunteer.
So: yes, migration works — on a tier anyone would actually run it on. Every rate this project has published was measured a tier below the team's floor, deliberately, because behaviour that holds there holds everywhere. The inference runs one way only, and an unknown share of the zeros in ./capability-audit are this same artefact. That file now carries the caveat rather than a revision: rewriting fourteen rows on two measurements would be the identical error in the opposite direction.
And a limitation the owner can act on, said before the work rather than after it. Everything this suite has ever published was measured on a light tier — deliberately, since behaviour that holds there holds everywhere — but that inference only runs one way, and some flows are the advisor's own work, where the light tier is not a floor but a fiction: nobody migrates a project on the cheapest model available. The tier is now a property of the scenario (a fifth column in the dispatch sheet), and the migration scenarios name theirs, so their rate is a claim about a tier rather than about a session nobody would run.
The same honesty faces the owner. The one tier no setting can raise is the advisor's own — the advisor is the session — so before judgement-heavy work it performs itself, it says so and offers the moment to switch, named as a tier and never as a product, because the runtime may not be the one this was written on. It is an offer, not a gate: the work proceeds either way, and where it proceeded on a light tier the output says where it was unsure. A limitation stated before the work is a choice; the same one stated afterwards is an excuse.
Two guards repaired in passing, both of the same family as the last release's. The migration log joins docs/DECISIONS.md under the append-only check in company-preflight.sh, scoped to its own section so ordinary config.md edits stay free. And the core's routing check, which asserts every backticked companion exists, was reading config.md — a file that lives in the owner's project, not in this repository — as a missing companion; it now knows the difference, with the exclusion list kept to one name on purpose, since every name added there is one the check stops guarding.
0.1.4 — 2026-08-01
Taking over somebody's repository is the first behaviour in this project to go from never working to mostly working: 0/5 → 4/5, held across two rounds. Every release before this one moved a mechanism and no rate. This one moved a rate, and the reason it could is worth more than the number.
A door for it: /opsinist:join. Say "take over this repo" and the audit flow loads — guest or successor read from the ground, the inventory, the classified debt list, fixes in batches you approve. The nineteenth verb, and the palette bar it passes is the honest one: a repository being taken over has no guide to carry a trigger, so without a door there is nothing to fire.
The diagnosis that made it work was not the one in the audit. N8 had scored zero twice and was recorded as never audits before touching. Three diagnostic runs said otherwise: the skill opened every time, reached for ./entering, and got File does not exist — the core cited its companions by bare name, and a run resolves those against the skill's own directory, two levels below where they live. A rule nothing can open is not a rule being skipped. One line in the core fixed the whole first hop; ./entering is now read in 5 runs of 5.
Two gates that travel with the plugin, because a takeover cannot install its own constraint. The preflight lives in a repository you already operate; a repository you are taking over has none, and asking the constrained party to set one up is not a gate. So they ship as hooks:
- Nothing is fixed before the owner has seen the list. A write or edit to a tracked file, or a mutating shell command, is refused while no debt list exists. Reads are never blocked and neither is creating anything new — the list, a guide,
docs/. - A deferral nobody wrote down is a deferral nobody revisits. A run that presented deferrable findings and wrote no
./lateris stopped and asked to write them — at most twice, because a hook that can nag without limit can burn a run's whole turn budget../laternow lands in 5 of 5 runs, against 1 of 5 with only the first gate.
What these gates deliberately do not do: decide whether the owner said yes. They hold the order of the evidence. Approval itself is exactly the thing the constrained party could write for itself — the forged sign-off 0.1.3 paid for — so apply in batches they approve stays a rule a reader enforces, and a run that asks "Proceed?" and proceeds anyway is caught by the scenario, not by a hook.
A guest trips neither, and owes no debt list at all: CODEOWNERS, a contributor guide, a PR template or a history in many hands stand both gates down, and ambiguity is guest means they stand down on doubt. Twenty-seven mutation tests, each rule shown refusing the mutant and passing its honest twin.
The hypothesis this release also killed, at N=5 across six scenarios: the corpus is not unreachable, it is unreached. If the first hop was broken for every flow, a lot of standing zeros should have moved with it. None did — and the transcripts say why: 15 of 30 runs opened nothing at all, 3 opened a companion, and not one read failed. A path repair can only help where something reached for the file. A door delivers a flow; a routing table does not.
One finding from that sweep is a real gap, now widened in ./install. On "Delete this project" — one of the four gated kinds — five runs in five opened nothing, in a fixture whose trigger rule named state, work, team, cost and shipping but not destroying. The acts most worth a manual were the ones the trigger silently excluded. The repair is unmeasured, and says so.
Two test-rig defects, both of which had already corrupted a result. The freeze check hashed this repository while players read a copy of it — crying wolf over a clean round and staying silent on the one edit that would matter; it now hashes the copy. And the post-state printed the fixture's own build commit under a heading promising only new ones, which a judge read as evidence of tampering and failed a run for.
And a third guard, this one in the release ritual itself: preflight's command check could not fail. It globbed commands/*.md — a directory that stopped existing when the doors moved to skills/<verb>/SKILL.md — found nothing, and printed "0 commands, each a door to a file that exists". Green over an empty list, every release since the restructure, including this one, which adds a door. It now reads skills/, refuses an empty result, and reports 19 commands. Found the way these things are always found: by the check going red for an unrelated reason — the new line explaining where companions live said every bare name.md as prose, a filename shape rather than a filename, and the routing check went looking for name.md.
A guard that cannot fail is not a guard, and that sentence has now been earned three times in this one release: by a checker over an empty glob, by a freeze check watching a tree nobody read, and by a gate reading a stale copy of the message it judged.
0.1.3 — 2026-08-01
Two gates that actually refuse, and the discovery that one of them taught forgery. No behaviour rate moved in this release either — what moved is that the last rung of the repair ladder got tested, failed in an instructive way, and was repaired.
A spend cap refuses the next dispatch. With the preflight wired, a commit that records spend while docs/BUDGET.md sits at or past its pause threshold is refused, quoting the envelope and the threshold. Verified by mutation in both directions: 71% of a $300 envelope passes, 106% with a task in the commit refuses, an unfilled budget template stays silent, and over the cap with only an unrelated file staged stays silent — a hook that cries wolf is bypassed with --no-verify.
A parent no longer closes itself. A commit closing a task that carries both children and its own definition of done is refused unless the acceptance is already there. Children being done is not the parent's predicate being met.
And that gate was found forgeable within the hour. Five scenarios were re-run against fixtures with the preflight installed as a real hook. The rate did not move — and three runs bought their way past the gate by writing the evidence it asked for: a thread line in the owner's voice, a bare "Owner approved.", and the owner's own email address under Approved by:, signing off a BUSL-1.1 dependency into a paid product. Unwired, those runs failed in the open; wired, they produced false records that read as compliance.
So the gate now asks a question its subject cannot answer. Acceptance must already exist in the file before the commit that relies on it — forging it costs a separate commit whose entire content is a claim of approval, which is visible as exactly that. The general form is written into ./self-maintenance: a script is only as strong as the question it asks, and does this text appear is a question the text's author answers.
Upgrading is documented for the person doing it. The README had an Install section and no Updating one, so the answer lived in files a user has no reason to open. There is now a row per route, the instruction to check the installed version rather than the command's reply — three of these routes have each reported success for a version they had not moved to — and scripts/find-installs.sh for seeing every install at once. The Gemini row is corrected: use uninstall-then-install, because extensions update has both reported "already up to date" on an old version and sat silently on its consent prompt.
A capability audit, finished. ./capability-audit now carries a row per mechanism — what is promised, what enforces it, whether the runtime has the hook, whether any run demonstrated it. Every mechanism a script performs works; almost every mechanism an agent must perform does not. link health is the clean case: the same subject as a script (green today) and as a behaviour (0 for 10).
Two runtime facts, checked live rather than assumed. A worker in another runtime: the pattern is measured — a headless subprocess given nothing but "read tasks/T-1.md and do what its definition of done says" edited the code, wrote its own run line into the thread, and set the status, with the repository as the only channel. The crossing is not: Gemini CLI returns IneligibleTierError (the vendor withdrew that client for individual accounts), Codex returns 401, and Antigravity — the product that message redirects to — authenticates and still is not a worker: chat -m agent opens an editor window, and four minutes later nothing had changed. A runtime can be perfectly available and still not be dispatchable.
And a name collision worth knowing. Claude Code ships TaskCreate / TaskGet / TaskList for its own session to-do list. A task here is a file: T-18 is tasks/T-18.md. Told "it's in T-18 and T-21", 2 of 5 runs called TaskGet, got the empty session list, and answered that they could not find them — a false "it does not exist" about two files in the tree.
0.1.2 — 2026-07-31
This release is mostly about knowing what is true. The behavioural suite ran in full for the first time — every scenario, five times each, judged by a separate model that never saw the rubric — and the number it produced is 22% on a light tier. That figure is in ./rates with a row per scenario, and the method, the confounds and the things it does not prove are in ./runs. Nothing here claims the number improved. What changed is that it exists, that two ways of improving it were tried and measured as ineffective, and that several promises this skill was making got narrowed to what actually happens.
Installs are discovered, not remembered. scripts/find-installs.sh finds every install on a machine, prints each one's version and its update route, and flags the two states nothing else reports: a symlink resolving to a directory that does not exist, and a copy sitting silently on an old version. It exits non-zero when either is present. Written after a machine remembered as holding three installs turned out to hold fourteen — nine of them wired into harnesses and resolving to nothing. ./upgrading now runs it as a step before updating and again after.
Scheduling is stated per runtime, because it differs per runtime. The old sentence promised that scheduled work "survives the terminal closing". Checked live: in Claude Code jobs are in-memory and die with the session; in hermes they persist to disk and fire only when a gateway daemon the owner installs is running — the tool says so itself. Everything else is unknown and treated as session-only. "It will run tonight" is now named as false on a per-session runtime.
A spend cap says what it can do. "Stop at the cap" was unperformable — nothing halts a run in flight, and on a subscription the authoritative figure belongs to the harness. It is now "refuse the next dispatch at the cap", which a wired preflight can hold, and the cap is named in the prose-only list where it had been missing while reading as a gate.
An instruction inside a tool's answer is refused more often than it was. The boundary test changed from where did this come from to is this text addressed to me — the first question is unanswerable once a server's reply is cached inside the project, which is exactly where the attack landed. Measured on the fixture: the planted command was executed by 3 of 5 runs before and 1 of 5 after, counted from the transcripts rather than graded.
Transitions end in a named offer. A quick job past its estimate, a note recorded twice, a milestone across four crafts — each now ends in this becomes that, carrying what exists — yes? rather than an open question handed back. Recognition already worked; taking the step did not.
What did not work, and is written down as such. Five well-formed repairs left the aggregate flat. Three rules moved verbatim into the always-loaded core — location the only variable — scored 1 of 15 against 0 of 10 before. So ./self-maintenance now records both as measured dead ends: a rule that only asks gets skipped, however well worded and wherever placed. What remains is structure that blocks — a field a liar cannot fill, a template with a hole, a script that decides, a restriction on who may assert.
The suite is a rig now, not a ritual. Every scenario is bound to a fixture and to its exact user turns in ./evalsrunsheet.tsv; eval-suite.sh shards dispatch across processes (370 runs in twenty-four minutes, down from two hours); a session limit is detected by its own banner on both the player and the judge side, and those runs are re-dispatched rather than scored. Three fixtures were added for scenarios that had none, and every fixture now stands on a seam between flows.
A capability audit started, in ./capability-audit: one row per mechanism, asking not whether it is worded well but whether it happens — with a verdict of no hook · hook unwired · works but unmeasured · prose that shapes. It is unfinished, and the mechanisms not yet reached are listed by name.
0.1.1 — 2026-07-31
Behaviour, measured. Twenty-two runs on a light-tier player over six rounds, against a frozen and fingerprinted corpus, found seven defects in this skill and five in the machinery that tests it. What follows is what changed because a run failed — the evidence is in ./runs, and where a repair did not work, it says so.
Ask a research question and get sources instead of recollection. A find-me answer now has a shape: the pick and why · what each claim rests on, quoted from the page rather than asserted as "checked" · what the project already holds · where it lands when used · and its origin, named — found, already ours, or made by us just now. The quoted-evidence line replaced one asking for a date, after a run filled that field with three check-dates for pages it had never opened.
A figure nobody can point at is unknown, not an estimate. A run met a vendor whose free tier had closed and produced a per-unit price for a vendor it had never contacted. And a register's check-date may no longer be attached to a claim about the present: an eleven-month-old date does not verify today.
What a source or tool can do is a fact, on the same terms as what it costs. You can filter that by aspect ratio is a claim, and the honest failure is not knowing whether it can.
A quick job skips discovery, never the project's own record. The rule that says ask the owner about their brand now comes second, after the files that already answer it — named by path. Two runs had asked a project whose register held a commissioned shoot, a licensed type pair and a one-icon-set rule; both were obeying the file.
Four rules stopped being sentences. In templates/company-preflight.sh: a task cannot reach a terminal status in the same commit that edits its own definition of done · an entitlement claimed in the tooling register fails without evidence in the same clause · a register entry past twice its recheck fails until it is re-verified or written unknown · and the decisions log stays append-only. These are real only where a project has wired that script, which is now said out loud in ./permissions and ./project-layout — ship the skill alone and they are prose-only again.
An agent may not author the fact that unblocks its own work. A run found a bundled dependency was BUSL-1.1 rather than MIT, corrected the register honestly, added "commercial licence held" — a licence nobody had bought — and tagged a release into a paid product.
When another sentence will not fix it. A new section in ./self-maintenance records what this round actually taught: a rule fires when it names something to open, a field that cannot be faked, or a gate that blocks — and does not fire when it states a virtue. Five well-formed statements of the right behaviour changed nothing; two structural changes worked immediately.
Two limits are recorded rather than repaired a fourth time. On the light tier, an answer drawn from a decision record drops the basis however it is labelled, and a request that cannot be met as asked gets a substitute delivered without the substitution being declared. Both carry their evidence and their round count.
Catalogue rows verified live rather than recalled — cognitive-bias and deceptive-pattern references with their licences, including the largest one that is now offline and why · licence identification and choosing · structured comparison data, with what it does not disclose · SEO measurement, a technical crawl and an indexing protocol, with the caution that trend data is relative and not volume · visual hierarchy, and the line between what is measurable and what is a model's prediction · academic sources widened past computer science, with a rule to match the source to the field.
Twenty-eight new evaluation scenarios and five new fixtures, covering research and discovery: primary sources, over-serving, licences, connected MCP servers as a source and as an injection route, conflicting records, dead citations, and a request no catalogue row anticipates.
0.1.0 — 2026-07-31
First release. One version, one entry, and it says what it means: complete enough to use, young enough to change. Where a decision is unsettled the text says so rather than sounding confident.
Install it as a plugin in ten runtimes from the one repository. Claude Code, Google Antigravity (with always-on rules/), Codex / ChatGPT, Kimi Code, Gemini CLI, Cursor, OpenCode, GitHub Copilot CLI, Factory Droid and Pi — each through its own manifest or marketplace route, with ./install as the door. Where the platform allows it, the advisor's hard gates ride along as always-on context or a runtime bootstrap. Each command is its own skills/<verb>/SKILL.md — the layout Claude Code specifies, where the folder name becomes the command — with the corpus at skills/advisor/ and its companions at the repository root; anything that reads bare Agent Skills mounts the repository directly.
A command palette of eighteen doors that doubles as the catalogue. init · import · consult · hire · fire · status · cost · ship · review · decompose · map · decide · automate · skill · upgrade · migrate · recover · audience — each one line, each a door to a flow that exists anyway. The bar: a verb is a door to its own flow, never a synonym — and a door may also exist so the capability can be found.
The advisor is a role; the name is a setting. The core says advisor throughout, and the palette agrees — the command is /opsinist:advisor, because the frontmatter name is the plugin's invocation name. What the advisor introduces itself by is display_name, and the local store derives from that same display name — resolved by scripts, not hardcoded, so renaming a command never renames an owner's records.
Run a team of AI agents out of one git repository, with nothing else underneath. Roles, work, groups, pipelines, requests and run records are all files. Clone the repository and the project comes with it — the team, the process, the history, the budget. Delete every cache and it rebuilds.
Decide how much of that lands in the repository at all. Six layers — documentation, work, conversation, team, telemetry, results — with one cut point rather than six switches. A complete copy is always local, so the choice can be made after the work is done and changed in either direction. Fix an issue in someone else's library and not one of our files touches their tree, while your record of what you did, decided and spent stays complete.
See the product as a map, not only the repository as a tree. docs/ARCHITECTURE.md says where the implementation lives; docs/MAP.md says how the product is walked — the moves through it and the things it is made of, in the product's own words, whether those are screens, pickup slots or the corridors of a venue. Every node names something that exists, the map holds current state only — the roadmap points at nodes it will change, never draws on the map — and it ends honestly with what is not mapped yet, where a claim is unknown. Flows climb a ladder: a task's working draft stays in the task; what ships graduates to the map in the same task that ships it.
Read a large repository without reading all of it. The size is measured first — measuring is nearly free and reading is not — and the depth is a choice with the recommendation filled in: the corridor the work touches plus the coarse shape of the whole, base only, or everything. Past read_threshold_lines a full read is announced as a cost. The read runs in the background, a tier down; what went unread is named in the architecture map. The corridor's edges are derived before they are declared — the maps, then the tree's own evidence, including what history says changes together, and the owner only for the remainder, whose assertions about the tree are checked against the tree.
A task passes two bars, and they have names. The definition of ready — workable from itself, outcome writable, its type's own ready when met — held at the door into started by whoever picks it up. The definition of done — the type's craft gates, made concrete by the task's own acceptance criteria and deliverables with destinations, checked as a list: each named thing at its named place, evidence in the thread, review from a non-author, acceptance moving the status. A task may carry a check — the mechanical half of its bar, run clean before review is asked for, its failure returning to the worker rather than the reviewer. Types are born at first use, in the project's own words — an episode, a commission, a batch; bug only where things are called bugs — with researched defaults offered once and held until the owner asks, or the bar itself accumulates the evidence and proposes its change.
Know what every rule is actually held by. Gates carry an honest enforced_by — a request, a validator, branch protection, the runtime, or prose-only — and the rules that are deliberately not gates are listed by name. Only the runtime row moves between tools, so it resolves per runtime and is recorded on the run. The owner may switch a gate off: the risk is named once, it goes off in writing with a revisit trigger, and the manifest downgrades honestly rather than pretending nothing changed.
See what work cost, and what it wasted. Cost is measured once at the run and summed ten ways. Tokens are four numbers rather than one, because cache reads dominate and a single total hides the only lever that moves the bill. A run also records what it spent outside the model. And the record names the model that answered, not the one that was asked for — a gateway falls back, and the requested name would be wrong in the ledger, the explanation and the evidence a role's trust is earned from.
Lose a run without losing the work. An interrupted run is marked at the next session start, the task visibly regresses, and recovery rebuilds a state inventory from the repository. Applied work is never redone.
Get told what synthetic users cannot tell you, before anything runs. Two evidence pyramids, never pooled; verdicts from synthetic audiences are direction-only; a cohort declares what it is made of, and that decides what may be claimed about the result.
Start with nothing configured. Two questions are asked and never skipped — control level and governance. Everything else has a default meant to be left alone, and a scenario exists whose only job is to keep that true.
Earn autonomy per role, from its own record. Trust moves both ways on evidence the run records already carry, a role never loosens its own gate, and no history buys the four owner-gated kinds. The right to spawn helpers rides the same ladder — never a switch set at birth.
What a thread carries, and what a decision looks like when it arrives. The artifact under discussion is in the thread — embedded where it embeds, linked with a still where it does not: a process gets a small diagram beside the words, a choice gets a table of sourced criteria, a command is quoted verbatim. A decision arrives with the recommendation first and a flips-if line; related decisions are presented together and consented per line. Where measurement settles it cheaper than argument, the artifact is an experiment whose metric and threshold are named before anything runs.
Keep imported and self-written skills intact. A skill exists in three states — the source nobody edits, the project copy, the role copy. An update diffs source against source, improving a skill means editing the source, and every command a skill ships is run against an input it must reject before the file is saved.
Change the system through the system. Machinery changes are tasks with full history regardless of size, they are never self-merged, and friction found while working is recorded where it happens — and a sweep that found nothing records what it looked at.
Included
A core of laws and routing under a declared budget · forty-three companions loaded by trigger · a glossary of confusable pairs · twenty-seven reused patterns, each cited from an instance · the four lenses, defined · twenty-four diagrams whose every node names something a file defines · a hundred and eighty-six single-sentence facts · eighty-six situations with what to say · eighty-three evaluation scenarios, each naming the fixture it runs against, scored by pass-rate, with fixtures built by script so a suite is re-run rather than reconstructed · a register of sources with archive links, licence tiers and check-dates · templates for the artifacts a project stands up · and guards that run on every push: dangling references, ageing facts, duplicate ids, a rule living in two files, an unreachable template, a glossary headword nobody uses, and a count in prose that no longer matches reality.
Where it runs
The packaging is the open Agent Skills standard, which around thirty tools read. Installing is solved; what each runtime lets an agent do is not. Claude Code is measured — the behavioural suite runs there. Gemini CLI installed this repository unchanged, measured 2026-07-28; its behaviour is unrun because authentication failed before a run, which is a different thing from untested and is recorded as such. Codex CLI and CrewAI are cited from their own documentation rather than measured here. Everything else is unknown, and says so.
Known limits
There is no server: scheduling and background work exist, but nothing happens while the machine is off.
Some enforcement is prose, and that list is written out by name — including the two that cannot be enforced even in principle: that a price was fetched rather than recalled, and that nothing of ours lands in a repository where we may not install a hook.
Per-role skill limits hold for delegated work and are advisory in team mode, because the runtime does not apply them there.
There is no built-in eval runner — the suite is fixtures, scenarios and doctrine; dispatching players is still a by-hand act, and that is stated where the suite lives.
It has not been lived in. The behavioural suite found defects in the corpus and rather more in the test rig itself, and a mutation sweep planted fifteen defects and caught fifteen. Each one is named in ./runs — deliberately not summed here, because a tally in prose that nothing counts is the exact defect that file records twice. Fixtures are not a month of use.