# `co audit` — is a CLI fit for an agent harness?

An agent that drives a command-line tool has only what the tool prints. `co
audit` asks one question of any CLI, `co` or anyone else's: can an agent with
only this program's help find the command for a task, run it safely, and know
what it will change?

Name the command as you would type it:

```bash
co audit co                  # all of co (about a minute)
co audit co gmail            # one co group
co audit co onenote          # OneNote help, including numbered pages and read
co audit yt-dlp              # any program on PATH
co audit gh pr --review      # then a model judges each page that passed
co audit co --json
```

It never reads source code. It runs `<command> --help` (or `-h`), opens every
subcommand the page lists, and keeps going down, many at once, each in an
empty HOME and working directory with no input and no API keys or tokens from
the environment. Then it judges each page from
what it printed. Exit 0 means fit; exit 1 lists each problem and its fix, and
a table scores every rule:

For a Python CLI installed with `pip --user`, it preserves the original Python
user-package path while isolating HOME. Otherwise the entrypoint could fail to
import its package before printing help, and the audit would mistake that for
an unreachable command.

```
gh: 228 pages
  prints         227/228
  hangs          228/228
  writes           1/228
  usage          228/228
  example        127/228
  ...
✗ not yet fit for an agent harness: 329 problems
```

## Rules (the same for every program)

| rule | fails when |
|---|---|
| `prints` | `--help` (or `-h`) does not print and exit 0 |
| `hangs` | it did not return within 20 seconds, e.g. waiting for input |
| `writes` | reading help created a file (the finding names it) |
| `usage` | no usage line |
| `example` | no example |
| `self_example` | no example runs this command itself |
| `flags` | an example uses a flag that the page of the command it runs does not document |
| `private` | an example contains a real home path or a full 0x address |
| `params` | an option or argument has no description |
| `look` | run again as a person's terminal runs it, the page says different words, or it is coloured in a pipe; for `co`, also a page with no colour in the terminal |

`look` runs each checked page a second time with `FORCE_COLOR=1`,
`TERM=xterm-256color` and 100 columns, with stdout and stderr each a
pseudo-terminal where the platform has one (the first run is `NO_COLOR=1`,
`TERM=dumb`, 200 columns, into pipes), so an audit takes about twice as long.
Both streams are terminals because a person's are: with stderr a pipe, the
`[env] …/keys.env` that opened every `co` command two or three times in a
real terminal never reached the audit (#2008). Words are
compared without colour, whitespace and frame characters, and in any order,
because a table narrowed to 100 columns folds a long value beside its
description; a value cut to `connec…` is a different word. A plain page in a
terminal fails only for `co`: many good CLIs print plain help on purpose, and
nothing about that stops an agent.

## Look (the standard `co` holds itself to, #1997)

Every `co` command prints through `connectonion/cli/style.py`:

- **Palette**: commands (examples, Next lines) bold cyan, counts bold, paths
  dim, warnings yellow, errors bold red, section titles bold underlined.
  A page written as one string goes through `style.markup()`.
- **Next line**: a result ends with `Next: <command>`, drawn by
  `style.next_line`. Only the command is coloured, not a parenthetical after it.
- **Colour by role only**: every console is made with `highlight=False`
  (`style.console()` is). Rich's default highlighter colours whatever looks
  like a number, date or path, so `co 1.9.0a5` came out with `1.9` alone in
  cyan and a date in three pieces; a word is one colour or none.
- **Titles are words**: no emoji in a panel title or a section heading
  (`📊 Account Status` became the section `Account`).
- **The first line is the command's own**: nothing printed at startup, and no
  line printed twice at the top.
- **Status layout**: a header line, then sections, one line per item; details
  behind `--verbose`.
- **Progress**: `style.progress()`, a bar with i/N when the total is known, a
  spinner with elapsed time otherwise; nothing is drawn off a terminal.
- **Plain when it should be plain**: no escape codes under `NO_COLOR`, in a
  pipe, a log or launchd, or with `--json`; the same words either way.

Beside the help pages, `co audit co` runs the read-only status commands in
`audit.STATUS` (`co status`, `co doctor`, `co commands`, `co rem status`,
`co auth status`, `co whatsapp check`) both ways, and fails a `Next:` line
there that is not drawn the way `style.next_line` draws it, or a status
command with no `Next:` line at all.

In a terminal run of any `co` page or status command, `look` also fails:

- the same line twice among the first five (startup noise such as `[env]`);
- in a status command's output, a word coloured in pieces: two runs of
  letters or digits inside one whitespace-delimited word in different styles,
  the mark of auto-highlighting (help pages are exempt: Typer colours
  `--options` and `<metavars>` on purpose, also inside a path);
- an emoji in a panel title.

Subcommands are read from the layouts real CLIs print: Typer/Rich panels
(any title but Options and Arguments), `Commands:`-style sections (uv, click,
kubectl), `CORE COMMANDS` with `name:` rows (gh), and argparse's
`{build,serve}`. A listed word whose page is its parent's page word for word
is not counted as a command.

## Model review (`--review`)

Only pages that pass every rule are reviewed. A text-only model judges five
things a rule cannot: is the first line clear to a newcomer, does the page say
what the command reads or changes, is the example realistic, is it simple,
and, when it acts on a listed item, can the user refer to that item briefly
instead of copying a long provider ID or full title? Commands that do not
select listed items pass the last question. It returns one concrete rewrite.
Pin `--model` when comparing runs.

## In CI

`tests/unit/test_cli_help_contract.py` runs the same engine on `co` on every
PR and blocks on any problem, `look` included. It adds three house conventions of ours: a
fixed "what it changes" word, a `Back:` line, and every command in
`co commands` reachable from `co --help`. The `help-gate` workflow runs
`co audit co --since base.json --review` on pages a PR changed, and reports
without blocking, because a model's verdict varies between runs. When the
model cannot run (no credit, provider error), the job summary says "review
unavailable" and the check carries a warning, so a missing verdict never
reads as a pass.

`co rem` keeps its own reviewed pages (#1656), which are being rewritten
(#1667); `co audit co` reports their missing examples. They are held to
`look` like every page, except the ones in the test's `LOOK_PENDING`, a set
that can only shrink and empties when their rewrite (#1996) lands.

OneNote also has a behavioral terminal journey in `tests/unit/test_onenote.py`:
`co onenote ls` → `co onenote pages 2` → `co onenote read 1`. The test checks
that both numbers resolve to the IDs that were displayed, even when two pages
share a title. The help review can flag a missing short reference, but it
cannot verify that a displayed number resolves to the right item.
