Loop Engineering
Stop accepting the first output. Design a process where AI iterates, evaluates, and improves — automatically.
Why this exercise?
Every prompt is a single roll of the dice. Even a great prompt. You send a message, the AI responds, and you either accept the result or manually re-prompt. Over and over. Most people accept mediocre output because the re-prompting cycle is exhausting.
Loop Engineering changes this. Instead of you iterating manually, you design a process where the AI:
- Runs — produces an output
- Checks the scorecard — evaluates the output against criteria you defined
- Retries or passes — if it fails the scorecard, it tries again. If it passes, it delivers.
The result: consistent, high-quality output with no babysitting.
"Stop writing better prompts. Design a process that retries until quality passes." — builders at the frontier of AI engineering
Your OCAREO Rules become the scorecard. The AI generates (creative, variable). The scorecard checks against your rules (fixed, consistent). When you separate these two, quality becomes repeatable instead of random.
Why one-shot fails
Here's a question: how many times do you re-prompt before accepting an output?
If the answer is "usually once or twice" — you're already doing loop engineering. Just slowly, manually, and inconsistently. You read the output, decide it's not quite right, adjust your prompt, and try again.
Now imagine a process that does this automatically. Not once or twice — but as many times as needed, against criteria you defined in advance.
| One-shot | Loop engineered | |
|---|---|---|
| Result | Single output, take it or leave it | Iterated output, only the best survives |
| Quality | Variable — sometimes great, sometimes off | Consistent — passes your scorecard every time |
| Quality bar | You, re-reading and re-prompting | Your 5 yes/no checks, applied automatically |
| Your role | The quality assurance | The scorecard designer |
Shopify's CEO ran an overnight loop on a single task — +19% improvement on a model that outperformed one twice its size. Later, the same approach made a rendering engine 53% faster. The principle: human owns the scorecard, AI owns the drafts. Keep only versions that pass.
The three-skill stack
Loop Engineering doesn't work alone. It's the third piece of a stack you've been building:
| Pillar | What it does | Your exercise |
|---|---|---|
| 01 — Spec | Define what "good" looks like (OCAREO) | TC14 + TC15 |
| 02 — Context | Fill gaps before generating — stop guessing | Step 0 below |
| 03 — Loop | Iterate until the output passes your scorecard | Steps 2-5 below |
They compose: OCAREO gives you the spec. Context-first fills the gaps. The loop runs until the scorecard passes. The better your pressure-testing (TC15), the better your loop performs.
Step 0 — Context before craft
Most AI failures aren't prompt problems — they're context problems. The AI doesn't have your company policy. It doesn't know last week's numbers. It doesn't know your client's preferences. So it invents them. And you don't notice until someone reads the output and says "that's not how we do it."
The fix: force a research pass before the AI generates anything.
Open ToqanClaw, paste your pressure-tested OCAREO spec from TC15, and add this block at the top:
Before generating anything:
1) List the top 5 pieces of context you are missing to do this well
2) For each one: can you get it from my connected tools
(Calendar, Drive, Slack) or do you need me to provide it?
3) Gather what you can; ask me only for what you can't find
4) Only then: generate against my scorecard
What this changes
Without context-first:
You: Write my weekly client update for Acme Corp.
AI: "This week, the team made good progress on the project.
Key milestones were achieved and we remain on track..."
→ Generic. No project names. No real risks. No numbers.
The AI invented a positive narrative because it had nothing else.
With context-first:
You: Write my weekly client update for Acme Corp.
[+ context-first block above]
AI: "I'm missing these 5 things:
1. Active project names — checking your Calendar... found 3
meetings this week: 'Acme Migration standup', 'Acme API
review', 'Acme QBR prep'
2. Last week's risks — checking Slack #acme-project...
found: 'API latency still above SLA' (Tuesday)
3. This week's deliverables — need you to provide
4. Client tone preference — need you to provide
5. Budget status — checking Drive... found Q2 tracker"
→ Real project names. Real risks from Slack. Real data.
It asked for 2 things. It found 3 on its own.
The difference: the AI researched before writing instead of guessing and hoping. Same task. Dramatically different output.
If your OCAREO Context section feels thin — if you wrote "my team" instead of "my 12-person product team across 3 time zones" — don't guess the details yourself. Let the AI research what it needs from your connected tools first. Context before craft.
Going deeper: Auto Research
Context-first works great for quick tasks with your connected tools. But what about problems in domains you don't know well — competitor analysis, regulatory changes, new markets?
That's where Auto Research comes in. Instead of one research pass, the AI runs a structured multi-phase investigation:
| Mode | When to use | What happens |
|---|---|---|
| Quick (context-first) | Your Calendar/Drive/Slack can answer | AI checks tools, fills 5 gaps, then generates |
| Deep (auto research) | Domain is unfamiliar or contested | Structured research sprint: explore → go deep → validate |
Try it: After your context-first pass, if you notice your OCAREO spec has assumptions you can't verify, ask:
I need to validate the assumptions in my OCAREO Context for
[your problem]. Research what you need — specifically:
1) What do I know for certain? (confirmed facts)
2) What did I assume that might be wrong? (contested claims)
3) What do I still need to answer as a human? (unknowns)
Don't rewrite my Objective or Rules unless evidence
contradicts them.
The AI researches, separates fact from assumption, and tells you what only you can answer. Your OCAREO goes from "confident guess" to "evidence-backed."
The Auto Research skill (available on the Files page) runs a full 5-phase research engine — explore, go deep, cross-validate, refine, and verify. Advanced, but powerful when you need to build a spec in a domain you're entering for the first time.
Step 1 — Install the skill
Download loop-engineering.zip — then install it in ToqanClaw: go to the Skills tab, click Upload skill, and select the .zip file.
Step 2 — See how OCAREO becomes a loop
Every letter in OCAREO maps to a role in the loop. This is the connection that makes the whole stack work:
| OCAREO | In the loop | Plain language |
|---|---|---|
| Objective | Definition of done | "What are we trying to ship?" |
| Context | What must be loaded each run | "What truth does it need before writing?" |
| Actions | The loop body | "Draft → check → fix → repeat" |
| Rules | Scorecard (YES/NO checks) | "What does 'good enough' mean?" |
| Examples | Golden standard / regression test | "What does a perfect output look like?" |
| Output | Exit format | "What artifact do we keep?" |
OCAREO is the brief you give a sharp junior. The loop is: they keep rewriting until you sign off. You don't rewrite for them — you own the sign-off criteria.
Step 3 — Build your scorecard
Before you design a loop, you need a scorecard — 5 yes/no checks that define "done." Pull out your pressure-tested OCAREO spec from TC15 and turn the Rules into checks.
Example: If your OCAREO Rules say...
| Rule | Scorecard check |
|---|---|
| NEVER auto-approve expenses above EUR 500 | Does the output flag high-value items for approval? YES/NO |
| ALWAYS accept receipts in any language | Does it handle non-English receipts correctly? YES/NO |
| ALWAYS categorise "home office" under "equipment" | Is categorisation consistent with our definitions? YES/NO |
| If a receipt is ambiguous, ask — do not guess | Does it ask for clarification instead of assuming? YES/NO |
| Flag alerts sent as DMs within 24h | Does the output include a notification workflow? YES/NO |
Your scorecard = your OCAREO Rules as yes/no questions. That's it. No code needed.
Step 4 — Design your first loop
Open a new conversation in ToqanClaw. Use your pressure-tested OCAREO spec from TC15 — or, if you didn't save one, try this:
I want to design a loop for the following task:
I need to generate weekly client update emails that summarise
project progress, flag risks, and propose next steps.
Currently I write these manually every Friday (takes ~45 min
per client, I have 6 clients = 4.5 hours). The quality varies
depending on how rushed I am.
Here's my scorecard (5 checks):
1. Does it mention all active projects by name? YES/NO
2. Are risks flagged with severity (high/medium/low)? YES/NO
3. Does every risk have a proposed mitigation? YES/NO
4. Is the tone professional but warm (not robotic)? YES/NO
5. Is it under 300 words? YES/NO
Design a loop that generates the email, checks it against
my scorecard, and iterates until all 5 checks pass.
Set a maximum of 5 attempts.
Step 5 — Test: one-shot vs. loop
Run the same task two ways:
- One-shot: Ask for the client email in a single prompt (no scorecard)
- Loop: Use the loop with your 5-check scorecard
Compare the outputs. The one-shot version will be decent. The loop version will be consistent — hitting every criterion you defined, every time.
The time math: 45 min x 6 clients = 4.5h manual. With a loop: 10 min review x 6 = 1h. That's 3.5 hours back every Friday.
Step 6 — Know the guardrails
Every loop needs safety rails. Without them, a loop that can never satisfy its criteria runs forever, burns tokens, and produces nothing.
Three guardrails every loop needs
1. Step limit — A hard maximum on attempts. "Try up to 5 times, then deliver the best attempt." Without this, a stubborn loop runs forever.
2. Cost ceiling — A maximum budget. "Stop if this costs more than X." Each iteration uses processing credits. Without a budget, costs compound.
3. Loop detection — "If you produce the same output twice, stop." Sometimes the AI gets stuck in a cycle — same input, same output, same failure. Loop detection catches this.
A real case: four AI agents ran for 264 hours. All health checks showed "running." Total cost: $47,000. Progress made: zero. The agents were running but not progressing. A green status light does not mean useful output — just like a weekly report that looks busy but says nothing. That's why you need a scorecard (progress signals), not just a status check (running signals).
When NOT to loop
Not everything should loop automatically:
- Legal advice — always needs a human gate before delivery
- Customer-facing content — loop the draft, but review before sending
- Financial decisions — loop the analysis, but human approves the action
- Anything with brand risk — the loop improves quality, the human owns the decision
The loop handles quality. You handle judgment.
Step 7 — Bring it to the Build Sprint
For the Build Sprint (Days 4-5), bring these three things:
1. Your pressure-tested OCAREO spec from TC15 — the specification
2. Your 5-item scorecard — the quality bar
3. Your guardrails — step limit + cost ceiling
That's everything you need to build something that runs reliably without babysitting.
OptionalGo deeper
- Stack all three: Start with a Pressure-Test session (TC15) to build the spec, then use the tested Rules as your loop scorecard. The better the test, the better the loop.
- The ownership model: You own the scorecard. AI owns the drafts. Only keep versions that pass. This is how the frontier works — not better prompts, better processes.
- Compound quality: Without scorecards, small errors in each step compound into big problems over a long workflow. Checking between steps catches issues early — before they snowball.
- For builders: The loop-engineering skill includes
auditor.py— a static analysis tool that checks Python loop code for missing guardrails. Advanced, but powerful if you're writing code.
What just happened
You learned the principle that separates one-shot users from people who ship things that work reliably: automated iteration with a scorecard. Humans iterate by instinct. Processes can iterate by design. When you combine this with OCAREO (for the spec) and pressure-testing (for the scorecard), you have a complete toolkit for producing consistent AI output.
"The difference is not a better prompt. It's a better process."
FAQ
Q: Do I need to be technical to use loop engineering? A: No. The concept — run, check the scorecard, retry — applies to any task. Emails, reports, analyses, presentations. The scorecard is just 5 yes/no questions. No code needed.
Q: How many iterations does a typical loop run? A: 2-5 for most tasks. The first attempt is usually 70-80% right. The loop catches the remaining 20-30% through scorecard evaluation. Rarely more than 10 iterations for well-specified tasks.
Q: What if the loop never passes? A: That's what the step limit is for. If the loop hits its maximum attempts without passing, it delivers the best attempt and tells you which checks failed. That's a signal to revisit your scorecard — the criteria might be too strict or contradictory.
Q: How does this relate to OCAREO (TC14) and Pressure-Test Your Spec (TC15)? A: OCAREO gives you the structure. Pressure-testing stress-tests the spec. Loop engineering uses the tested Rules as the scorecard for automated iteration. They're three stages of the same pipeline: specify, validate, execute.
