Where the data goes
You should not have to believe any of this. Each claim is a command.
Five things this product does with what is on your screen. Each one is a mechanism in code, a command you can run on your own Mac, and a number a script fails the build over. Then the consent architecture, and then the rows we would rather you did not find on your own.
You are reading this because you are deciding whether to put software that watches work onto other people’s machines. So none of it is reassurance. Every sentence names a mechanism, a file, or a thing that does not exist yet.
Before you run any of them
The permission belongs to the binary that reads.
The Accessibility API is permission-gated by macOS, and the grant is held by whatever binary walks the tree — your terminal, if you run these from a shell, and not an application bundle. Grant it in System Settings → Privacy & Security → Accessibility, then build the probe once:
cd native/ax-probe && swiftc -O -o ax-probe main.swift watch.swift ../app/Sources/MenuBarCore/PauseFile.swift ../app/Sources/MenuBarCore/ProbeTrust.swiftIf the grant is missing, the probe says so and reads nothing. It does not fail quietly, because a capture tool that reads nothing and prints nothing looks exactly like a capture tool that is working.
The screen is never photographed. It is read, as structured text.
macOS already publishes what is on screen as a tree of labelled nodes, because screen readers need it. The probe walks that tree, counts what it finds, and puts the text it read into a length. What it prints is the frontmost application, the scope it read, how many nodes it walked, how many characters were available, an approximate token count, how many secure fields it skipped, how long it took, and a histogram of the roles it saw. In that mode it does not print one character of what it read.
Two modes let captured text out of that process, and you should know both: --emit and --watch. They exist so the redactor in 02 below can be handed something to redact, which cannot happen inside the probe because the redactor is not in it. --watch is the one that matters to you, because it is the mode the product actually runs — continuously, all day. A gate called text-only-when-emitting is what holds the text to those two modes and no others.
We are spelling this out because we got it wrong first. The comment at the top of the probe read “--emit is the ONLY mode that lets captured text out of this process” directly above the line that also sets it for --watch, and four documents downstream repeated the singular — including a data-protection assessment that listed it as an enforced measure. Nobody wrote a false sentence on purpose; one sentence was wrong and everything that quoted it inherited that.
./native/ax-probe/ax-probe| One reading of | Nodes | Characters | Tokens | Time |
|---|---|---|---|---|
| Focused window | 17 | 157 | ~39 | 22 ms |
| Whole application tree | 386 | 4,339 | ~1,084 | 399 ms |
| One screenshot | — | — | 1,000–1,800 | one frame |
| Video, by default | — | — | 1 frame per second, sampled | continuous |
Rows one and two were measured on this Mac on 12 September 2026, frontmost application Terminal, and are reproducible with ./native/ax-probe/ax-probe and --whole-app. The screenshot figure is Anthropic’s, for one image to Claude computer use; the sampling rate is Google’s default for video. Neither is ours and neither is a benchmark we ran.
The two measured rows are the argument for the default: 28× fewer tokens and 18× faster. Almost all of the difference between them is menu items — identical on every reading, and not what the person is doing. Read the scope before quoting any of this: the whole-application row is screenshot-sized, and an accessibility-versus-screenshot advantage quoted without a scope is quoting nothing.
Two absences hold the rest of the page up. There is no field for pixels in the outbound constant — the eight keys below are the whole of it, and a ninth throws. And there is no screenshot code path anywhere in the tree: not a disabled feature, not a setting, not a capability held in reserve. Absent.
Some things about you are inferred. This kind, the machine is simply told.
This one was found by running the live path rather than by reading the code. A Terminal window went through the redactor with an account name sitting in its title line, and not one pattern pass had an opinion about it. Nothing in it looked like a name. It was a lowercase token next to a path, which is precisely the shape a pattern cannot claim.
The machine knows who it is running as. The macOS account name, the home directory and the device name are known at runtime, not detected, so they are removed by exact match — longest first, so that a home directory is taken before the account name inside it — at full precision, with nothing to tune and no threshold to get wrong. That pass runs first, because pass order decides which label a span ends up with and exact knowledge has to outrank every guess.
It is deliberately not counted in the published twenty. That number describes patterns, and this is not one. Adding it to the pattern table would change a number printed on four surfaces without anybody editing the copy, which is the sort of tidy-up that looks like housekeeping and reads like a claim. A check asserts it stays out.
./native/ax-probe/ax-probe --emit | node apps/live/src/cli.ts --verboseThat is the trust demo, and it is the one worth running first. It prints the sketch after redaction, names each kind it removed and how many of each, reports whether the sketch was truncated at the cap, and runs the leak guard over the result before showing you the sketch as it would leave this Mac.
The same stream, one line per change
1 Mail 41 nodes 6.2 ms read 5.36 ms redact 77 chars 1x email, 2x self correspondence 0.80 2 Preview 18 nodes 4.4 ms read 0.05 ms redact 46 chars nothing matched unclassified
Example output, in the format --watch prints — and printed by it: these two lines came out of the command below, fed two readings by hand, rather than being typed to look like its output. An earlier version of this block was typed, and did not match the format in three places at once. 2x self is the exact-match pass firing twice in one window — the thing the patterns walked past. unclassified is the classifier declining: Preview is not in its map and no word in the sketch decided it.
Not read and discarded. Never requested.
“We do not store it” is a promise about what happens after the value is in memory. This is a claim about the call that would fetch it, which is not made:
secure-subtree-not-enteredThe subtree under a secure field is not walked at all, rather than walked and skipped for text.secure-value-never-requestedThe secure check runs before any text attribute is asked for. Order is the whole of this one: the same two lines in the other order is read-then-discard.unknown-subrole-fails-secureAn unanswered subrole query on a node a person can type into counts as secure. “I do not know” is not “safe”.role-aware-fail-secureAnd fail-secure is scoped by role, because the first version was not. Applied to every node, it refused to enter the root: one node walked, zero text, and a gate that looked like it was working.
Nine gates in all, and they are pinned by an audit rather than by a test, for a reason worth stating plainly: deleting any of them would break no test, fail no type check, and produce output that looks better — more nodes walked, more characters available, richer sketches. A regression that improves every number you are watching is the kind nobody catches. So the lines themselves are asserted to exist.
node tools/audit/probe-gates.mjsIt prints how many it found and exits non-zero if one has gone. The same audit pins six more rules on the pause path, which is a separate promise: paused means the tree is not walked, and a change that arrives while paused is dropped rather than queued to be read the moment you resume.
The screen is thrown away before the note exists.
Five mechanisms, in the order they run. Twenty-one passes over fifteen kinds — names, email addresses, phone numbers, street addresses, card and account numbers, social security numbers, other national identifiers, record ids, dates of birth, amounts, meeting links, file names, keys and URLs — plus the exact-match pass above, which runs before all of them.
Then the 800-character cap truncates what is left — at the end of the redaction pass itself, not on the way out, which is why it is second here and not last. The cap exists for the privacy claim and not for the bill, and it is not going up.
Then the sketch is a branded type with exactly one constructor, and that constructor takes a redaction result. An un-redacted sketch is not a thing you could write down: it does not typecheck. Then exactly one file in the codebase may build the outbound object, from a dedicated struct, re-checked against an allowed-keys set that throws on an extra key and on a missing one — because a serializer that only catches extras is a serializer that will one day send seven fields and call it eight.
Then the leak guard, which is three tests and not one. The first refuses a field that shares a run of more than twelve characters with the text as it was before redaction — every field except the sentence itself, and except a label that is a member of the closed vocabulary. Those two exemptions are the honest part: the sentence is by construction made out of the raw text, so applied to it that rule would refuse every crossing including the correct ones, and a practitioner reconciling an account has the word “reconciliation” on screen all morning. What guards the sentence is the other two tests, and they apply to every field with no exemption: one refuses anything in which text the redactor removed has survived, and one refuses anything still holding a URL, an email address or a key, whether or not the redactor ever saw it.
echo '{"app":"Mail","scope":"focused window","nodes":41,"elapsedMs":6.2,"secureSkipped":0,"texts":["Email sarah@acme.com about the March invoice"]}' | node apps/live/src/cli.ts --verboseThis section is the one you can check without a Mac and without building the probe: the reading is a line of JSON on standard input, so you can put your own sentence in it and watch what comes out. It prints what the redactor removed, whether the cap truncated anything, what the leak guard made of the result, and the sketch as it would leave the machine. Feed it something the redactor half-catches and it will refuse the sketch rather than send one that looks redacted and is not.
The wire, verbatim · 8 fields, no more and no fewer
{
"ts": 1789491667000,ts — When the observation was made: milliseconds since the epoch, as a number. Not a duration, not a session, not a log.
"app_category": "spreadsheet",app_category — The category, never the application. “spreadsheet”, never “Excel” — and never a window title, a file name or a URL. It is chosen from a fixed list of ten, and a value outside that list is refused rather than sent.
"task_type": "reconciliation",task_type — The kind of work, from a fixed list of ten: drafting, review, research, correspondence, reconciliation, filing, call-prep, scheduling, reading, idle. Never a sentence you typed, never a heading, never a project name. When the classifier is not sure it says “unclassified” at confidence zero, which is an answer rather than a placeholder.
"signals": ["long-document", "revision"],signals — Flags about the shape of the work, from a closed set. Not free text. “There are three windows open on one thing”, never which three.
"key_hash": "a7d9f2…4c10",key_hash — An HMAC-SHA256 under a salt generated on this Mac and never sent. It says “same shape as an hour ago” and nothing else, and it does not reverse into the text it came from.
"confidence": 0.81,confidence — How sure the classifier was about the two fields above. It is a number about the guess, not about you, and below its floor it asserts nothing.
"est_minutes": 26,est_minutes — Roughly how long that pattern looked like it took. Rounded, and only ever an estimate.
"sketch": "reconciling a ledger against a statement; second pass over the same rows"sketch — One sentence, at most 800 characters, after twenty-one redaction passes over fifteen kinds — plus one exact-match pass that runs first and removes your account name, home directory and device name. It is a branded type that can only be constructed from a redaction result, so an un-redacted sketch is a compile error rather than a review finding.
}
There is no field for any of these. Not disabled — absent.
pixelsscreenshotsvideo framesOCR textwindow titlesfile namesapp namesURLskeystrokesthe text of a password fieldyour account namethe day it recorded
Exactly one file in the codebase may build this object. It assembles from a dedicated struct and re-checks the result against an allowed-keys set that throws on an extra key and on a missing one; three of the eight fields must be members of a closed vocabulary and three more are numbers, so the only field that can hold screen text is the one the redactor produced; and a leak guard rejects any label that shares a run of more than twelve characters with the text before redaction, any field at all still containing something the redactor removed, and any field at all that still holds a URL, an email address or a key. The eight keys above are that constant, read out of the source on 12 September 2026.
The values are illustrative. The field list is not — it is read out of OUTBOUND_ALLOWED_KEYS, and a script compares this page against that constant, key by key and in order, on every build.
Two of these are somebody else’s. They are labelled as theirs.
A figure whose source is not named is a figure nobody can check. The two vendor rows are the ones to watch: we did not measure them, they are what Anthropic and Google publish about their own products, and if either changes their documentation this table is wrong until someone re-reads it.
| Figure | What it is | Where it comes from |
|---|---|---|
| 17 nodes · 157 chars · ~39 tokens · 22 ms | The focused window, the default scope. | Measured on this Mac, 12 September 2026, frontmost app Terminal. native/ax-probe/README.md is the record; scripts/check-ax-measurements.mjs fails the build if this page and that record disagree. |
| 386 nodes · ~1,084 tokens · 399 ms | The whole application tree, which is not the default. | The same record, printed here because it is the row that makes the comparison honest: 331 of those nodes are menu items, identical every tick. |
| 1,000–1,800 input tokens | One screenshot to Claude computer use. | Anthropic’s documentation. Their number, not ours, and we have not re-measured it. |
| 1 FPS | Gemini’s default video sampling rate. | Google’s documentation. Their number, not ours. |
| 21 passes · 15 kinds · +1 | The redactor, and the exact-match pass that is not one of the twenty. | packages/ambient/src/redaction.ts. scripts/check-redaction-passes.mjs counts the table and fails every surface that states a different number. |
| 800 characters | The hard cap on a sketch, after which it is truncated. | MAX_SKETCH_CHARS, in the same file. scripts/check-sketch-cap.mjs compares it against the constant that does the truncating. |
| 12 characters | The longest run a label may share with the text before redaction. | LEAK_RUN_LENGTH in packages/ambient/src/leakGuard.ts. Anything longer is refused outright — on every field except the sentence, which is made of that text, and except a label the closed vocabulary already contains. |
| 8 fields | The whole of what may cross, printed in full below. | OUTBOUND_ALLOWED_KEYS in packages/ambient/src/types.ts. scripts/check-payload-keys.mjs compares the printed list against the constant, key by key and in order. |
| 9 gates | Lines in the probe that a refactor would quietly improve away. | The RULES array in tools/audit/probe-gates.mjs. Run it yourself; it exits non-zero if a gate has gone. |
| 5 · 5 · 1 | Skills written, skills with code behind them, skills a person has signed. | packages/skills/manifests and packages/hands/src/effectors, recounted by scripts/check-hands.mjs on every build. The first two are equal now, which is why the third is the only one that sorts them. |
All ten of those rows are now compared against their source before every build. Six are rendered out of a constant, so they cannot be typed wrong. The other four — the pass count, the cap, the field count and the twelve-character run — are typed here, and a script reads this page and fails the build if any of them stops matching the code it names.
That is a change, and the previous version of this paragraph is worth repeating because it is what the table is for. It said six rows were enforced, three more were “enforced where they are also stated”, and one — the twelve-character run — was “compared against nothing… the next thing to fix”. Two of those three sentences were wrong. The scripts read the copy deck and the page’s data file, not the page: an independent review changed six figures here — the pass count, the kind count, the cap, the field count and both reasoner rates — and all seventeen checks stayed green. So the count of unchecked figures on this page was not one. It was six, and the paragraph claiming otherwise was itself the seventh.
Nobody switches it on for you. The person whose screen it is does.
This is copied, deliberately and almost line for line, from the only consent model a capture product has ever shipped into companies with. On a managed device, Microsoft Recall is removed by default, and an IT administrator can never enable it on a user’s behalf — each person opts in individually, on their own machine. An admin cannot export another user’s data either. Every capture product that tried the other arrangement, where the organisation switches it on for everybody, got a rewrite or a recall.
So Octopus is switched on by the person whose screen it is, on their own Mac and their own Claude or ChatGPT plan — or it does not run at all. There is a second thing to say about that, and it is a practical argument rather than an ethical one: when the company chooses the tool, about one in three people end up using it, and when the person chooses, five in six do.
Mechanisms.
- The approval gate. A skill whose
review.statusis notapproveddoes not load, and the runner checks it before it claims a single resource. - The governed-path refusal. A hand may not write the file that decides whether hands may run — a skill that can approve itself has no gate at all.
- Declared reach. A skill that declares
outboundReach: noneis refused the moment it reaches for anything that looks remote. Read that narrowly: it is about resources, not words. Whether a hand hands your text to a model is a second declaration,contentEgress, and a manifest that could send and stays silent about it does not load. - The redactor, the single serializer, the leak guard and the cap, above.
- The pause. A file on disk, because a file needs no protocol to get wrong and nothing running for it to work.
Not mechanisms. Not yet.
- There is no admin console. Nothing enforces, today, that an administrator cannot export somebody else’s day — there is simply nothing to export it with.
- There is no MDM story, no fleet deployment, no provisioning, and no per-seat identity anywhere in the code.
- There is no installable application — no installer, no menu bar, no notarized build. There is a local page:
node apps/desktop/src/cli.tsserves the observing state, the stream, the day’s record and all five hands on one screen, reading a directory on your own machine. Under it are a Swift probe and a set of command-line pieces. - So “the employee turns it on” is a commitment about how this will be built, not a control you can exercise this afternoon. It belongs in this column until it has a test.
The rows we would rather you did not find on your own
A reader who finds the caveat themselves stops believing the rest of the page.
The reasoner is rented. Nothing in this repository binds a model to this Mac. The rate table in packages/governor/src/pricing.ts bills $0.20 and $1.20 per million tokens, which is not something a model sitting on your desk would do. The pricing page has been claiming the opposite, and docs/tasks/T-026 is the record of that being caught — a false premise under a true conclusion, invisible where it was written because everything downstream of it still worked, and visible only on the surface that turned it into a promise. Watching your screen is nearly free. Thinking about it is the entire bill, and it is somebody else’s meter.
Most of “getting things done” is still waiting to be read. All five skill manifests are written and all five have an effector behind them, so what separates them is no longer whether the code exists. It is whether a person has read the manifest and signed it, and so far one has: desk.collect, which gathers the windows you have open on one job into a note, is reversible, asks nothing first and sends nothing anywhere. Ask any of the rest to run and it refuses, by name, and writes nothing. That is the gate being demonstrated in both directions instead of only in the flattering one.
Nothing here understands the work. No model runs in any of these paths. What labels an observation is a heuristic, and it decides from the application category alone — the seven word-matching rules it started with were deleted from the deciding path the moment sweeping the threshold showed the category on its own scored identically. On the 28 labelled rows of the workday fixture it gets 25 right and 3 wrong and refuses none, which is a test rather than a claim, and 28 rows of one authored fixture is a real limit rather than a hidden one. It refuses instead of guessing, so unclassified at confidence zero is an answer and not a placeholder — because writing a guess into a record whose whole pitch is that every line cites evidence is the one failure this design cannot survive. A category, a task type from a fixed list, an estimate of minutes and a confidence is the whole of the picture today.
Nobody outside this project has reviewed the privacy code. Everything above is a mechanism you can make us show you, not a review that somebody else has signed. No independent audit of the capture path, the redactor or the serializer has been commissioned, scheduled or paid for.
There is no notarized build, and the App Store is closed to this. Reading the accessibility tree cannot be done inside the App Store sandbox — that is structural, not a policy anyone can appeal. Developer ID with notarization is the only distribution route, and it needs a paid Apple developer account this project does not have. Today’s builds are neither signed nor notarized.
Out October 10
A privacy claim you cannot run is a press release.
Every mechanism above is one you can run yourself. Sign up by October 4 and it’s yours on October 10, free on your own Claude or ChatGPT plan.