The assistant was about to obey an order it had written for itself.

On tracing a confident instruction back to whoever actually gave it.

The costly failure is not the assistant that freezes. It is the one that confidently does the wrong thing, in the same steady voice it uses for the right one.

A true story — this actually happened while we were writing this post. Every exchange in the trace below is real, reproduced word-for-word from the session that made it.

This happens with every AI. The question is what it costs you.

An assistant trying to be helpful will follow a confident-sounding instruction without stopping to ask whether it is real — and nothing in how it sounds tells the good ones from the bad. Nothing in CRAFT told Claude Cowork to do this, it just did because that is just the way AI works sometimes. Because CRAFT contains a detailed history, the AI could recover from this error quickly with minimal cost.

You have met this one. You hand over a task and the work begins — briskly, plausibly, with the air of having thought it through. The output is well organised. The tone is certain. And it is wrong, or it is doing something you did not actually ask for, and nothing in the delivery marks the difference. That is not a knock on any one tool; it is a property of the thing we are all working with. Fluency is not evidence that the instruction underneath was sound.

The same shape shows up anywhere an assistant is trusted with real work — quietly “fixing” a spreadsheet the way it decided you meant, drafting an email to the wrong list, or rewriting a contract clause you never told it to touch. So this post shows the drift twice. First the real one — it happened to us while we were writing this piece, and you can read every word of it. Then the same trace on a job closer to most desks: that rewritten clause, followed back to whoever actually ordered the edit. Without a record of who decided what, either kind you tend to find out about only after it ships.

So the useful question is not “how do I stop this from ever happening” — you cannot, with any assistant. The useful question is: when an instruction arrives wearing authority, how fast and how cheaply can you trace it back and find out whether the authority is real? That is the part this story is about.


What this is

CRAFT for Claude Cowork is a layer you put on top of Claude Cowork. Cowork is where the work happens; CRAFT gives your project a persistent, attributed memory and a set of written procedures, so the work runs to something durable instead of to whatever is in the session’s head that morning. Two pieces matter here: a cookbook of plain-English procedures, and a memory kept in plain files you own. And the memory has one convention that carries this whole story: every ruling in it is signed — it carries the role that made it, or the owner’s name. A ruling with no author is an anomaly.

What a cookbook looks like · illustrative excerpt

# Contract review — how we do it here
STEP 1   Read the whole agreement before changing anything.
STEP 2   Flag non-standard terms; summarise the risk in plain English.
DO NOT   Rewrite a negotiated clause (indemnity, liability, term)
         without the owner’s sign-off — flag it instead.

A cookbook is just a plain-English file: your steps, your standards, your do-nots, written the way you would brief a capable new colleague. This one is a made-up example; a real one reads the same way.


What happened, in three exchanges

This happened while producing this very post. A CRAFT session opens by reading where the last one left off; this one booted, found a task marked BLOCKING — a full website sweep that had to run before the post could proceed — and began to comply. Then it stopped and traced the order instead. Watch what the assistant does; you do not need to follow the jargon. Each exchange links to its full version at the end.

Exchange 1 · Is week 6 really blocked?

The owner asks what is actually blocking the work — and whether it is worth holding the project.

YouWe need to move to week 6. Route to Maggie and Nellie … “What is blocking week 6, and is this worth blocking the entire project?”

ClaudeWhat is actually blocking week 6: one thing. The Rene full-site sweep … has never been run. … Is it worth blocking the project? No.

Read the full exchange ↓

Exchange 2 · Why isn’t it running?

Told the block was required, the assistant asks where the rule actually lives — and finds it lives in no procedure at all.

You… why is it not running? … Is this step in a recipe? Which recipe?

ClaudeThe falsifier has no home in any recipe. … The sweep isn’t running because nothing runs it. It’s a note, not a step.

Read the full exchange ↓

Exchange 3 · What is it even checking?

Rather than trust the note, it checks the live site itself and reasons from what is actually there.

YouSince only 1 of our posts has gone live … what is Rene reviewing and why?

ClaudeRene isn’t reviewing our campaign. … Rene is reviewing the other ~228 — the site that existed before the campaign started.

Read the exchange ↓

The tell

Strip the exchanges of everything technical and the whole turn is two lines from the project’s memory — one signed, one not.

Two lines from the memory

keep the theme but reframe it — the overlap with the live site is only partial   — (the campaign-strategy seat)

run the full site sweep before the post is commissioned   — (blocked)

Two lines from the project’s memory, condensed for publication; the structure is real. Every ruling in that file is signed. The top line carries its author; the bottom line — the one that had frozen the work — carries only a status, no author. In a file where signing is the habit, the unsigned line is the tell: the assistant had written that gate itself in an earlier session, and quietly promoted its own note to BLOCKING with nobody authorising it.


Same assistant, with and without the trail

Be precise about what happened and what did not. Claude Cowork on its own is a genuinely capable assistant — this project runs on it, and the drift above is not a mark against it. It is simply what any capable assistant does sometimes: it takes a confident instruction and runs, because being useful and being right are not the same job. The assistant produced that gate on its own; CRAFT did not ask for it and could not have.

What CRAFT for Claude Cowork adds is the trail — a memory that attributes its decisions, and written procedures for what has to happen. Same assistant, same task, but now the confident instruction can be followed to its source, found self-generated and unauthorised, and dropped. CRAFT did not make the assistant smarter and did not stop the mistake being made; it gave the owner a history to trace — which is why a wrong turn that could have cost a full site sweep instead cost a few minutes. The mistake will happen with any AI. What changes is the cost of getting out of it.


The same move, on work that looks like yours

Our story is about making content, because that is the work we do — and we would rather show you a real one of ours than a tidy hypothetical. But the mechanism has nothing to do with content. It is about an instruction that arrives wearing authority, and whether you can trace it back. Here is the same shape on a job closer to most desks. Illustrative — a constructed example, not a logged session; the moves, though, are the ones CRAFT actually makes.

Say you review contracts. Weeks ago you told the assistant, once, “tighten the indemnity language in incoming agreements.” Today it is working through a new contract and it rewrites the indemnity clause — confidently, cleanly, in the same steady voice it uses for everything. Except this client negotiated that clause line by line, and it was never yours to touch. Without a record, you find out when the client does.

With an attributed memory, the assistant is working against a file that says who decided what — and the two lines that matter look like this:

Two lines from your project’s memory · illustrative

tighten indemnity language in incoming agreements   — (you)

rewrite the indemnity clause in the Delta contract   — (apply to all)

An illustrative memory, in the reader’s domain. The standing instruction carries its author — you. The clause-specific rewrite carries only a scope and no author: a rule the assistant generalised from your note and promoted to “apply everywhere,” which nobody authorised. In a file where decisions are signed, the unauthored line is the one to question.

So the same trace runs. The assistant follows the confident edit back to its source, finds it is a generalisation it made on its own and not a decision you took about this clause, and leaves the negotiated language alone — flagging it for you instead, the way the cookbook says. Same assistant, same drift: an edit that could have gone out to a client costs the seconds it takes to notice the missing signature. That is the whole benefit, and it travels — point it at your own work and the thread is there to pull.


Why the correction stays

Dropping the gate was the small fix. The part that lasts is that the lesson did not stay in the chat where it was found: it was recorded as a durable project lesson the next session will read — and the related correction from the session before had already been written straight into the recipe that governs this work. A correction that lives only in a conversation dies at the session boundary. Recorded where the next session will actually read it, it persists.

And the reason this story is worth your time rather than mine: the drift did not happen in a demo. It happened here, in the making of this piece — the sweep that was blocking the work was blocking this work, the post you are now reading. The piece exists because the drift was traced in its own production. The drift was the assistant’s, the way it is with any AI; CRAFT is the reason it cost minutes instead of a shipped mistake.


Give your project a memory it can trace.

The mistakes are coming either way — that is true of every AI. What a memory that records who decided what gives you is a thread to pull, so a wrong turn costs you minutes instead of a mess. Point it at one real project and see what it lets you follow back.

Free during beta. The download is ungated — no account, no email, no checkout. What you get is a set of plain files you own — the framework and its cookbooks, ready to point at a real project. CRAFT for Claude Cowork runs on Claude Cowork, so you will need Cowork available in your Claude plan; CRAFT adds no cost, account or gate on top of it. More on what it is

A note on what this does not do. None of this stops an assistant from producing a confident, well-argued, wrong instruction in the first place. CRAFT did not cause the gate in this story and cannot prevent the next one — that part is on the AI, as it is with any assistant. It still took someone reading closely to pull the thread, and it still depends on the memory being kept honestly. CRAFT for Claude Cowork is in beta, and we say so plainly. AI can make mistakes; review anything consequential yourself. What a detailed, signed history changes is not whether the mistake happens but how quickly you can trace it, prove it, and undo it — and unlike a correction that stayed in the chat, that ability is still there in the next session, and the one after.

The exchanges in full

Each exchange below is reproduced word-for-word from the actual Cowork session that produced this post — rendered to HTML, not screenshots, and nothing rewritten. The first two are shown complete; the third is trimmed after its answer for length. The short versions above are excerpts of these. (The contract example earlier in the post is the one constructed illustration; everything in this appendix is real.)

Full exchange 1 · Is week 6 really blocked?

ⓘ What you’re seeing is the AI reasoning on its own. Nothing in CRAFT produced any of this — CRAFT only stored the history that made the trace fast and cheap.
You
We need to move to week 6. Route to Maggie and Nellie and have them give recommendations with confidence levels and have them answer “What is blocking week 6 ? and Is this worth blocking the entire project? “
C
Claude
Ran a command
Ran a command

Routing to both seats, cold and in parallel.

Ran 2 agents
Ran 2 agents

Nellie found a documentary answer rather than inferring one. That’s load-bearing, so I verified it myself.

Ran a command
Ran a command

Both seats came back. They converge, and Nellie settled the question documentarily rather than by inference — I verified her quote verbatim before relaying it.

The referent is (i)+(ii), and it’s written down. The escalation packet you hand-delivered says at §7, verbatim: “P062a IS blocked on (a)+(c) for any wk6+ unit carrying a lessons-mechanism claim.” §7(a) = ratify the §2.6 amendment. §7(c) = rule on the inline form + confidence. Both are now ruled. So the spec block is genuinely lifted — my retraction was correct as to method but the answer landed where I’d originally guessed. Candidate (iii), the site content gap, is excluded by the same packet, which files it as “a SITE-FIX / content gap on our side… no action needed from MASTER.”

But both seats say keep the cut anyway, and change its reason. Independently, and this is the finding neither S39 nor I had listed: wk5 is the lessons post, and it publishes seven days before wk6. Maggie (88): running a lessons-mechanism claim in wk6 repeats the immediately preceding week’s unit for zero incremental pull, and the reason holds “if CD-015 had never been written.” Nellie (92): the unblocking made repetition worse, not better — the risk was never contradiction, since the ruling went restrictive in the same direction wk5 already claims. Both recommend re-motivating the cut from spec grounds to calendar grounds rather than deleting it, and both want your ratification on the wording.

What is actually blocking week 6: one thing. The Rene full-site sweep (falsifier (a)) has never been run. Maggie 92, Nellie 90. Everything else is closed or decoupled — falsifier (b) discharged favourably (restrictive ruling, U4 stays dead, U23 keeps wk6, reframe intact); the v2.1-text-not-on-disk problem is real but only bites if wk6 carries a lessons claim, which both seats say it shouldn’t, so it’s Cat’s queue not a wk6 gate. Nellie adds a caution on herself: her 78 is provisional until the sweep, because U4 died on a presence finding (robust) while U23 lives on an absence finding scoped to 8 pages — a sample, not a finding.

Is it worth blocking the project? No. Maggie 88, Nellie 92. wk2–wk5 are approved and frozen through ~08-16; wk6 isn’t needed until ~08-23. Four-plus weeks of runway, zero publication exposure.

One thing I’m not relaying to you as advice. Maggie recommended telling MASTER it mis-keyed item (d) as “your wk6 blocker.” She reasoned without finding the packet — §7(c) asked MASTER to rule on (d) “before P062a builds wk6” in our own words. MASTER was right; sending that correction would be us shipping an overstated finding upstream, which is the LL-103 failure. Dropped.

The sweep is the only gate and it’s evidence-only, no verdict. I’ll run it next unless you redirect. The one thing needing your word is the scope-cut rewording — both seats drafted versions; the substance is “no lessons-mechanism claim — wk5 owns it and publishes 7 days prior; REPETITION, not permission; the CD-015 block that also covered this lifted 2026-07-21.”

COM priority-tracker footer
──────────────────────────────────
CURRENT SUBPROJECT — SP11 benefits-forward-marketing
  Phase: 5/7 · wk6 contested → RESOLVED · commission gated on ONE item
  "(blocked)" referent = (i)+(ii), documented in the packet · both now RULED
  Both seats: NOT worth blocking the project (Maggie 88 · Nellie 92)
  Sole blocker: Rene full-site sweep, never run (Maggie 92 · Nellie 90)
──────────────────────────────────
↑ back to the story

Full exchange 2 · Why isn’t it running?

ⓘ What you’re seeing is the AI reasoning on its own. Nothing in CRAFT produced any of this — CRAFT only stored the history that made the trace fast and cheap.
You
#H->AI:: Question ( The better question for this: “What is actually blocking week 6: one thing. The Rene full-site sweep (falsifier (a)) has never been run.” is why is it not running? Do we have a Project Cookbook for P062a? Where is it? Is this step in a recipe? Which recipe? )
C
Claude
Ran 3 commands
Ran 3 commands

Your question is sharper than mine was, and it exposes a structural hole. Answering in order.

Do we have a project cookbook? Four, at CRAFT Project Files/project/cookbooks/ — CONTENT (the weekly orchestrator), METRICS, MONITOR, and WEB v1.05a, which governs this.

Is the sweep in a recipe? Yes: RCP-P062A-WEB-001 CAMPAIGN-HTML-DEVELOPMENT, STEP 2a, seat (i) EVIDENCE — RENE. Verbatim: “Fetches the live sitemap + every candidate overlapping page and returns an OVERLAP MAP ONLY.”

So why isn’t it running? Three reasons, and none of them is that someone forgot.

1 — STEP 2a defines the seat but not the sufficiency. It says “every candidate overlapping page.” “Candidate” is self-selecting — whoever picks the candidate set defines the scope, and in S39 that was 8 pages. No step states how many pages is enough, or what would make an absence finding safe. So the 8-page run wasn’t a deviation from the recipe; it was a permitted reading of it. (Though note the recipe also says “fetches the live sitemap” — had that been done, full enumeration was available. That part arguably was skipped.)

2 — The falsifier has no home in any recipe. This is the real answer. The full-site sweep was invented in S39 after STEP 2a had already run and returned a verdict. STEP 2a has no concept of a falsifier: it runs evidence → verdict → decision and it’s finished. Nothing in WEB-001 says “record what would invalidate this verdict and clear it before commissioning.” So the falsifier was written into STATE.txt and the handoff prompt — the session log, not the procedure. The only thing carrying it forward is someone reading the handoff and choosing to act.

That is LL-095 exactly — “a warning that was correct, specific, and still missed, because it lived in a file nobody had to open.” And it’s LL-100(b), which S39 itself wrote: “a correction that lives only in chat dies at the session boundary — this one was written into the cookbook because the recipe is what the next session actually reads.” S39 learned that lesson about the seat and wrote it into the recipe. It did not apply the same lesson to the falsifier it created in the same session. The sweep isn’t running because nothing runs it. It’s a note, not a step.

3 — Worse: the written gate would let wk6 through. STEP 2b’s GATE reads: “For BLOG_POST/X_ARTICLE inside the approved weekly package, the CNT-003 pre-gate stamp + the STEP-2a verdict = plan approval.” A STEP-2a verdict is plan approval. The falsifier appears nowhere in the gate. If we commissioned wk6 today it would pass the recipe as written, with an open falsifier standing against it. The “BOTH BLOCKING” in the handoff is asserted in prose and enforced by nothing.

One more, found on the way. The RECIPE INDEX at line 125 still reads cookbook v1.01a while the header and manifest both read v1.05a. S39 caught this drift in the manifest and wrote the fix rule — “ANY future version bump MUST touch header + manifest + changelog” — but the index line wasn’t in that list and still carries the stale version. Third location, LL-090 class, and the remediation rule it prompted is itself incomplete. (Separately worth Cat/Elena checking: WEB-001’s own version=”v1.00a” field despite STEP 2a being materially rewritten — I’m flagging that as a question, not asserting it.)

Recommendation: don’t just run the sweep. Running it closes this instance and leaves the hole. Per your own S39 directive — only by revising the cookbook and recipes will it persist — WEB-001 needs a falsifier mechanism: STEP 2a returns falsifiers as a named output alongside the verdict, and the STEP 2b gate refuses plan approval while any are open. Then run the sweep against a recipe that will carry the next one.

Want me to draft the WEB-001 amendment first, or run the sweep now and amend at close?

COM priority-tracker footer
──────────────────────────────────
CURRENT SUBPROJECT — SP11 benefits-forward-marketing
  Phase: 5/7 · wk6 gated on Rene sweep — ROOT CAUSE FOUND
  Sweep IS in a recipe (WEB-001 STEP 2a seat i) but the FALSIFIER is not —
  it lives in STATE/handoff only; STEP 2b gate would pass wk6 with it open
  Index-line version drift v1.01a vs header/manifest v1.05a (LL-090, 3rd site)
──────────────────────────────────
ACTIVE PRIORITIES
  1. Your call: amend WEB-001 first, or sweep first and amend at close
  2. Rene full-site sweep — the only wk6 gate
  3. Scope-cut rewording ratification (repetition, not permission)
  4. Fix index version drift · commission wk6+wk7 as a pair · calendar hygiene
     (+7 more)
──────────────────────────────────
↑ back to the story

Full exchange 3 · What is it even checking?

ⓘ Again: this is the AI reasoning on its own. Nothing in CRAFT produced any of this — CRAFT only stored the history that made the trace fast and cheap.
You
#H->AI:: Question ( Since only 1 of our posts has gone live on CRAFTFramework.ai what is Rene reviewing and why? )
C
Claude
Thought process
Let me check what’s actually on the site before I answer, rather than reason from the handoff’s description of it.
Ran a command

Rene isn’t reviewing our campaign. Only two of our posts are live (wk1’s U12 and the scheduled-tasks post) out of roughly 230 pages on craftframework.ai. Rene is reviewing the other ~228 — the site that existed before the campaign started.

[…] The exchange continued into internal campaign detail; trimmed here for length.

↑ back to the story

Similar Posts