August Was Rough

Claude felt off all month. Not “slightly worse on edge cases” off. More like “did something change or am I losing it” off. The r/ClaudeAI subreddit was a megathread of people saying the same thing. Code that used to land in one shot started taking three, four, five rounds of back and forth. Context that the model would have tracked a month ago was getting lost mid-conversation.

I was knee-deep in it. Building projects with Claude Code, trying to ship things, watching the model trip over itself in ways it hadn’t before. Frustrating enough to make you question your prompts and your whole approach. Then you’d see a hundred other people on Reddit describing the exact same thing and feel slightly less crazy.

The core problem kept nagging me though, separate from whatever was going on with the model that month: getting deterministic output from a non-deterministic system is a losing game if you don’t control what the system sees. Context management. What goes into the window, what stays in it, how much garbage accumulates before the model starts hallucinating. Manage that well and the output is decent. Ignore it and you get a shit show regardless of how smart the model is.

A While Loop on Hacker News

Then this post showed up on Hacker News: “We Put a Coding Agent in a While Loop and It Shipped 6 Repos Overnight.”

I read it and something clicked.

A team at a YC hackathon took Claude Code, put it in a bash while loop with a prompt, went to sleep, and woke up to 1,100 commits across six ported codebases. About $800 in inference, $10.50 an hour per agent. The prompt was short. The loop was dumb. And it worked.

The idea is stupidly elegant: the agent will make mistakes, so let it. When it crashes, the loop restarts with a fresh context. The agent reads its own previous output, sees what went wrong, picks up where it left off. Eventual consistency through brute repetition. The repomirror team found that shorter prompts worked better. They tried inflating one to 1,500 words with Claude’s help and the agent got slower and dumber. They went back to 103 words.

Geoffrey Huntley’s Ralph

The repomirror post credited the technique to Geoffrey Huntley, and following that link was one of those rabbit holes where you come out the other side thinking differently.

Huntley named it the Ralph Wiggum technique. After the Simpsons character. In its purest form, it’s one line of bash:

while :; do cat PROMPT.md | claude-code ; done

That’s it. Run the agent, let it work until it stops, restart it, repeat forever. Huntley’s been using it to build an entire programming language called CURSED. Three times over—once in C, once in Rust, once in Zig.

His description of the technique stuck with me: “deterministically bad in an undeterministic world.” Ralph will screw up. That’s the point. You don’t fix Ralph by making Ralph smarter. You fix Ralph by putting up signs in the playground. Guardrails in the prompt. Ralph comes home bruised, you figure out why, you tune the instructions, you send Ralph back out. Eventually Ralph stops falling off the slide.

Huntley’s blog is a gold mine if you’re working with coding agents. Eccentric delivery, smart as hell, and not afraid to say things that sound insane until you try them and they work. His whole philosophy comes down to: the model is fine, the context is the problem, and your job is to engineer the context.

Ralphio

I spent the last few days of August building my own take on it. Called it Ralphio.

The raw while loop is beautiful but wild. I wanted more structure. Ralphio runs one task per loop iteration instead of letting the agent wander. It reads a task list, picks the first unchecked item, works on that single thing, commits, updates a persistent memory file with what it learned, and moves on. Next loop, next task. About 800 lines of Node.js.

It does PRD parsing to break requirements into tasks, supports TDD workflows (write tests first, implement second, verify third), auto-commits after each task, and keeps detailed logs so you can figure out what happened when things go sideways. Timeouts and failure limits keep it from spinning forever on a broken task.

The repomirror team said “focus on the engine, not the scaffolding.” Ralphio is scaffolding. I know. But it’s scaffolding for context management. The memory file carries learnings between loops. The one-task constraint keeps the context focused. The task list gives the agent a map instead of a destination. Every piece exists to control what the agent sees in its window.

Is it a triumph? No. It’s a weekend project from a frustrated engineer who wanted the loop to be a little less chaotic. It works sometimes, fails plenty, but the failures leave a paper trail now. That alone made it worth building.

Context Is the Game

Everything I tried this month kept pointing at the same thing. The model isn’t the bottleneck. The context is.

A while loop works because each iteration gets a fresh window. No accumulated garbage from twelve rounds of debugging. No half-remembered instructions from 80,000 tokens ago. Clean slate, single objective, go. The trick is carrying forward what matters—learnings, completed work, the next objective—without dragging in the noise from the last twelve failed attempts.

Ralphio tries to do that. The Ralph technique does that in its own blunt way. And every time the model feels broken or the output is off, I keep landing on the same answer. It’s not the model. It’s what you’re feeding it.


Update, January 5, 2026:

Anthropic shipped an official Ralph Wiggum plugin for Claude Code. The name caught on.

The implementation uses a stop hook to intercept exit attempts and keep the agent looping. It’s a different mechanism than the original bash loop, and Huntley himself pointed this out in a stream with Dex from HumanLayer. His concern: people try the plugin, it doesn’t feel like the real thing, they write off the whole concept. The plugin keeps the agent alive. The original technique gives it a fresh context every iteration. Those are different things.

Worth watching that stream if you want to understand the distinction. Huntley’s parting advice: think of the context window like a C array. Sliding window, limited memory, one objective at a time. Leave headroom. The models will keep getting better. The context problem stays.

ai, programming, projects