A skill workflow ran for an hour. It had spawned its own tmux session somewhere, issued a command, and the terminal hung. The agent waited for output that was never going to come. I had no way to know any of this. The process said “running.”

I killed it. Started over.

I built Claude Code Watcher earlier this year for exactly this reason. Headless Claude Code with the -p flag gives you no terminal, no output, nothing to look at while it works. The Watcher fixed that for Claude Code. But I stopped using Claude Code as my main agent.

The Blind Spot Grows

I switched to Pi a while back. I’ve got skills running headless now that spawn sub-processes, manage their own state in JSON files on disk, chain into other skills. I wrote about the prompts-as-code paradigm that makes all of this work. The short version: markdown files are the programs, JSON is the state, the agent executes against them.

The more I leaned into this, the less I could see. A skill takes 20 minutes. Working or stuck? A workflow makes 50 tool calls. Are half of them redundant? One of my skills got into a loop where it retried the same failed bash command four times before giving up, and I only found out because I happened to read the transcript after.

Claude Code Watcher didn’t help. It only watched Claude Code.

Claude Code Watcher → Agent Peeper

So I added Pi support. I also added filesystem polling for Claude Code. The hooks from the original Watcher still work, including in -p mode, but the filesystem scanner doesn’t need any setup. It checks ~/.claude/projects/ for JSONL files modified in the last five minutes. No settings.json config, no monitor.js to install.

Pi was easier. Pi writes its transcripts to ~/.pi/agent/sessions/. The app polls for recent files, reads the session header, done. No hooks to install, no configuration.

The two agents use different transcript formats though. Pi has a tree structure with id/parentId for branching conversations. Claude Code is flat JSONL. Agent Peeper normalizes everything to one format so there’s one rendering path. Tabs show up with badges: CC in blue, π in violet.

At that point it wasn’t a Claude Code tool anymore, so I renamed it.

How It Works

Three discovery mechanisms run in parallel, all feeding the same dashboard:

Agent Peeper architecture diagram showing three discovery mechanisms: CC Hooks, CC Filesystem Polling, and Pi Filesystem Polling, all feeding into a live Electron dashboard
How sessions get discovered.

It’s an Electron app. Reads files from disk, renders them in a dashboard. No server, no database. Sessions expire after five minutes of inactivity. Polls every two seconds for new transcript entries, newest first.

One HTML file for the entire UI. No React, no bundler. I keep the whole thing as dumb as possible on purpose.

Agent Profiling

The stuck process was what made me build it. What keeps me opening it is the waste. A skill reads the same config file three times because context degrades between steps. A workflow shells out to jq ten times when one query would do. I watched an agent spend four minutes generating a response, throw it out, and start over because the prompt was underspecified.

It’s profiling. Same idea as profiling code, except the program is a markdown file and the runtime is an LLM. You watch, spot the hot paths, rewrite the prompt. The model itself is still a black box. Nobody, including the people who built it, knows what the fuck is happening inside. But the workflow around the model is a markdown file and a JSON state file. Those I can actually work with.

This Gets Worse

OpenClaw has 200K+ GitHub stars. You give it a Mac Mini, a messaging app, and a model. It runs 24/7. You text it a task and it goes off and does… whatever it does. Sometimes for an hour. Maybe you get a message back when it’s done. Maybe you don’t. What happened in between? No idea.

I’m working on something in that direction too. Less “AI with its own social media” and more “AI with a computer and an email address.” More on that soon. The observability problem gets worse as agents get more autonomous. Not better.

Agent Peeper is on GitHub. MIT license.

ai, programming, tools, projects