Case Study · Personal Project

App Doc Generator

An agent that learns an Android app by actually tapping through it, then writes the help-center article itself, with screenshots and demo videos. I came up with the idea and the logic behind how it works, then built it by directing AI coding agents.

Built with AI coding agents Next.js + TypeScript Python · Appium Model Context Protocol SQLite (FTS5) ffmpeg pipeline

How this was built

I didn't write this code by hand. I built the whole tool with AI coding agents. My job was the thinking: what the tool should do, how each part should work, and what to change when the results weren't good enough. The agents turned those decisions into code.

My part
  • The idea: document an app by actually using it
  • The approach, like reading the screen's structure instead of guessing from pictures
  • The writing rules that make the docs read like a technical writer wrote them
  • Testing real runs, spotting what was wrong, and deciding how to fix it
The AI's part
  • Writing the code for each part I described
  • Wiring up the libraries (phone control, video, database, web app)
  • Making the changes I asked for after each test

The sections below go into technical detail. Read them as the decisions I made and why. The code that carries those decisions out was written by AI, following my direction.

App Doc Generator dashboard showing projects, published docs, recent runs with engine and cost, and connected devices
The home dashboard. Every row under "Recent activity" is a real documentation run, tagged with the engine that drove it and what it cost.

Most "AI documentation" tools summarize a document you already wrote. This one starts from a running app instead. You give it a feature name and a sentence of instructions. It opens an Android emulator, finds the feature on its own, performs it for real, records a clean demo clip as it goes, and comes back with a formatted how-to article ready to publish.

I built it because I was tired of watching documentation rot. Every time an app shipped an update, some screenshot somewhere was suddenly wrong, and nobody noticed until a user complained. The project started as a question: can an agent drive a real app reliably enough that I'd trust what it writes? It turned into a full web app I actually use.

The problem with documenting an app by hand

Writing a how-to article for a mobile feature is three jobs stacked on top of each other. First you perform the flow yourself, carefully, so you don't miss a step. Then you write it up in plain, consistent language. That's more tedious than it sounds, because "tap the field and enter your company name" needs to become "fill in the required details", or the doc is wrong the moment someone's account has different custom fields. Then, if you want a demo video (and you should, because screenshots alone are a worse experience), you do the whole thing again on camera. Fat-finger one tap near the end and you re-record the clip from scratch.

None of those jobs is hard on its own. Doing all three, correctly, for every feature, every release, is what doesn't scale. That's why documentation is usually the first thing to fall behind.

What it actually does

From the outside, using it looks like this:

1
Point it at an appPick an app installed on a running emulator, or on a real phone over USB.
2
Describe one featureA name and a sentence, like "Document creating a new lead, starting from the home screen." No screen-by-screen script.
3
Pick what drives itYour own API key, or an existing Claude Code or GitHub Copilot subscription. Same behavior either way.
4
Watch it happen liveA live panel streams every tap, typed field and screenshot over a WebSocket as it happens. You can see exactly what it's doing and cancel at any point.
5
Get back a real documentIntroduction, prerequisites, one section per action it found, numbered steps, a cover image and short demo clips, already trimmed and placed in a phone frame.
6
Review and shipAnnotate a screenshot, re-cut a video by hand if the auto-trim guessed wrong, then export a self-contained HTML page or a Markdown-plus-media zip.
New project wizard: choose a device, then an installed app
Add a feature form with document types: how-to guide, UI overview, revamp and video capture
Left: a project starts by picking a device and an installed app. Right: each feature is a name, a doc type and optional instructions. The four doc types cover a fresh how-to, a screen overview, rebuilding an outdated doc, and adding videos to a doc someone already wrote.

Architecture at a glance

There are two halves. A Node server owns everything the user sees: projects, prompts, the knowledge base, the video pipeline. A small Python server owns the phone. Between them sits one of three interchangeable agent "engines", and the only thing any engine talks to on the phone side is a fixed set of MCP tools.

Scroll sideways to see the full diagram →

Browser · Next.js UI Dashboard · live run view · doc editor · screenshot annotator · video trim editor HTTP + WebSocket (live run events) Node server · Next.js + ws (server.ts) Run orchestrator runner.ts · one run at a time · caps Prompt builder prompt.ts · style guides · writing rules Project store SQLite · runs · doc versions · media Knowledge base SQLite FTS5 search over help articles injected into prompts + served at /api/kb/search Video pipeline content-anchored trim · crop · phone frame ffmpeg · QC status · manual trim editor starts one engine per run LLM API Anthropic or OpenAI-compatible ApiRunner TypeScript agent loop bring your own key Claude Code CLI runs on a subscription GitHub Copilot CLI runs on a subscription MCP over stdio · same tools for every engine MCP server · app-doc-driver (Python) tap · type_text · swipe · press_back · screenshot · begin_action / end_action · discard_step · search_kb · done accessibility tree → numbered element list · status-bar demo mode · settled-frame anchors for trimming search_kb (HTTP) raw clips + timing sidecars WebDriver Appium + adb UI tree XML · gestures · screen recording Android emulator or USB phone any installed app, no source code needed
Highlighted boxes are the pieces I'd call the core: the orchestrator, the in-house agent loop, and the MCP server. The dashed lines are the two side channels the phone side uses to reach back into the Node server.

A run goes top to bottom. The orchestrator builds one prompt (feature, instructions, style guide, knowledge-base snippets, writing rules) and hands it to whichever engine the user picked. The engine loops: read the screen, call one tool, read the result, repeat until it calls done(). Screenshots and raw clips land in a per-run folder. After the run, the video pipeline turns the clips into finished videos, and the Markdown gets rendered with its media markers resolved and saved as a new doc version, so older versions can always be restored.

The core idea: read the tree, not the pixels

A lot of "AI controls an app" demos show a vision model a screenshot and ask it to guess pixel coordinates to tap. It looks great in a GIF and is unreliable in practice. Screen density varies across devices, a button moves eight pixels between versions, and estimating the exact center of a small icon from a picture is something vision models are still shaky at.

So this doesn't do that. Android exposes an accessibility tree for every screen, the same data screen readers use, and Appium can hand it over as XML. The tool parses it, drops everything that isn't visible, tappable or carrying meaningful text, and compacts what's left into a short numbered list:

[ 12] Button        "Save"          id=action_save   bounds=(900,80,1060,160) clickable
[ 27] EditText      "Company"       id=company_input bounds=(80,320,1000,400) input
[ 31] ImageView     ""              id=fab_add       bounds=(940,1780,1060,1900) clickable

A trimmed example of the compact tree the agent actually reasons over.

The model just says "tap 12". No coordinate math, no guessed pixels. Every gesture tool (tap, type, swipe) returns the refreshed tree in its own response, so the agent never spends a turn re-reading the screen after acting. IDs only belong to the tree they came from, so a stale ID can't quietly tap the wrong thing three screens later.

It isn't pixel-free everywhere. Some Jetpack Compose icons render with no label, description or resource ID at all, so there's a coordinate fallback for those. The prompt is explicit that it's a last resort, and in a normal app the tree covers almost everything.

One tool surface, three brains

The phone is controlled through a small local server that exposes tap, type_text, swipe, press_back, screenshot, begin_action/end_action for recording, and done as tools over the Model Context Protocol, on top of Appium. That server doesn't know or care which model is calling it.

There are three ways to drive it. The first is a TypeScript agent loop built for this project that talks straight to an Anthropic or OpenAI-compatible API with native tool calling, so anyone with an API key can use it. The other two spawn the Claude Code CLI or the GitHub Copilot CLI, which already speak MCP and bill through whatever subscription the person already has.

Settings page with three AI engine options: API, Claude Code CLI and Copilot CLI, plus emulator and Appium controls
Runs tab for the Stopwatch feature listing two API runs and one Claude run with their costs
Left: the engine is a setting, and can also be switched per run. Right: the same Stopwatch feature documented three times, twice through the API engine and once through Claude Code, with the cost of each run.

That was a deliberate call. I didn't want the project to live or die by one vendor's agent harness or pricing. The prompt and the tools are identical whichever engine is driving; only the outer loop changes. It also turned out to be a useful debugging trick. Running the same instructions through all three shows you exactly where the models make different judgment calls, which you'd never see building against just one.

The tapping turned out to be the easy part

Getting the agent to find and perform a feature took real iteration, but it came together faster than I expected. Getting it to write documentation that reads like a technical writer published it, rather than a transcript of button-mashing, took a lot longer. Most of that work lives in the prompt, not the code.

The clearest example is what I ended up calling the universal-steps problem. To record a good demo, the agent has to fill in real forms with realistic data: every visible field, a real-sounding name, a plausible address. Left alone, it then writes that up literally. Tap Company, type the company name, tap the back arrow, enter First Name, enter Last Name, tap Lead Source, choose an option. It's a faithful record of what happened, and it's useless as documentation, because the next person's account has different custom fields and the steps are wrong on day one.

What it wants to write

Tap the + button.
Tap Add Lead.
Tap the Company field.
Type the company name.
Tap the back arrow to confirm.
Enter the First Name.
Enter the Last Name.
Enter the remaining fields.
Tap Lead Source.
Choose an option.
Tap Save.

What ships

Tap the + button.
Tap Add Lead.
Fill in the required details.
Tap Save.

The fix is a rule, not a smarter model. Filling a form is one step in the text, full stop. The field-by-field detail belongs in the video, where it helps, not in prose that has to stay accurate across accounts and releases. A handful of rules like that did more for the output than any amount of asking for "concise": steps capped at twelve words, one action per step, exact button labels in bold, and never narrating "you will see...".

A generated how-to article for a feature called Quick Mode, with introduction, prerequisites, a list of possible actions, short imperative steps and a note
A document the agent wrote for "Quick Mode" in an AI character app: a short introduction, the actions it discovered, short imperative steps with bold labels, and a note for the one real caveat. The phone screenshots and clips are left out of this capture.

There's a second layer. A lot of mobile features are configured on the web and only used on the phone. A small keyword-search index over existing help articles is queried before a run starts, and the agent can search it mid-run too. That lets it tell "this is new on mobile, describe exactly what I see" apart from "this is set up on the web, so say so and link to it". The prompt is blunt that the live screen always wins if the two disagree.

The video pipeline was the hardest engineering problem

The idea is simple: the authoring run is the recording. The agent wraps each action in begin_action() and end_action(), which start and stop a screen recording on the emulator, and whatever it does in between becomes raw footage. There's no separate "now do it again for the camera" pass.

The catch is that raw footage is a mess. It includes the model's thinking pauses between tool calls and any wrong taps it recovered from, and it runs far longer than the action takes. The obvious fix is to timestamp each tap relative to when recording started and cut the clip to those times. I tried that first. It doesn't work, because the recorder's clock and the actual video timeline drift apart (encoder warm-up, variable frame rate). Trust the clock and you cut into the previous action, or freeze on a half-loaded spinner.

What works is anchoring by content instead of by clock. After each gesture settles, the tool server saves a screenshot of the settled screen plus a small fingerprint of its UI tree. When the raw clip is processed later, image matching finds the frame in the video where that settled screen actually appears, and the keep-window is built around that frame rather than a guessed time. Clock drift can't affect it, because it never looks at the clock; it looks at what's on screen. When nothing matches cleanly (an old run with no settled shots, or genuinely ambiguous frames), it falls back to the time-based cut and flags the clip for review instead of quietly shipping a bad one.

Raw clip
Trimmed
kept, anchored to a real settled frame cut: thinking gaps, retries, dead air

A few more layers sit on top. discard_step lets the agent mark its own mistakes, a wrong tap and the Back press used to undo it, so they're removed from the video entirely. An optional vision-model check looks at each clip's last frame and asks whether the screen has actually finished loading, and extends the clip if it hasn't. Every clip also carries a status of auto_ok, needs_review or edited, because I don't trust any single heuristic, mine included, to be right every time. For the cases it gets wrong, the web app has a non-destructive timeline editor: scrub, split, drag a boundary, save, and it re-renders from your edits without ever touching the raw clip.

Small details that mattered more than I expected

Some problems only showed up once I was watching real footage, and each needed its own fix.

The rest of the app

Under the agent it's a real web app, not a script with a UI bolted on. Next.js runs behind a custom server so WebSocket and HTTP traffic share one process. Projects and features are SQLite rows, and every generated document is saved as a new version you can roll back to. The live run view streams every thought, tool call and screenshot as it happens, and asks "still going?" if a run runs long, so nothing burns through a budget unnoticed. There's a Fabric.js screenshot annotator for redacting or highlighting before publishing, per-project style guides that set tone and feed reference examples into the same prompt every engine gets, and export to a single self-contained HTML file or a Markdown-plus-media zip, depending on where the doc needs to live.

What I'd change

A few things are still rough, and I'd rather say so than pretend otherwise.

The clock-drift problem is really a symptom of the recorder I chose

Content-anchored trimming works around the emulator's built-in recorder drifting from real time. There's an experimental host-side recorder (scrcpy) that gives noticeably smoother footage, but it only emits frames when the screen changes, which breaks the anchoring assumptions in a different way. It's still in there, opt-in and off by default, and honestly unsolved.

The three engines don't explore identically

Same prompt, same tools, but different models make different calls about how thoroughly to explore before deciding a feature is covered. I haven't closed that gap, and I'm not sure it can be closed completely without giving up the point of supporting more than one engine.

Error handling leans on "log it and keep going" more than I'd like

It's built to run unattended for twenty minutes or more, so a lot of non-critical steps (a failed screenshot, a demo-mode command that didn't stick) are best-effort rather than fatal. That's the right call for keeping a long run alive, but it also means some failures are quieter than I'd want in something other people depend on.


What building this taught me

Going in, I assumed the hard part would be the agent: getting a model to navigate a UI it had never seen. That came together faster than I expected, mostly because reading the accessibility tree instead of guessing at pixels removes most of the fragility up front. What actually took the time was everything after "the agent found the right screen": writing prose that stays true a release later, and turning a recording full of hesitations and mistakes into something that looks intentional. Neither of those is a model-capability problem. They're editorial and engineering problems that happen to have an LLM in the middle, and treating them that way, with explicit rules, fallbacks and a human review step instead of blind trust in the automation, is what made the output usable. It's also how I worked with the coding agents that built it: they wrote the code, and I judged every result against what the tool actually had to do.