Give it a job.
Open the finished file.
Safata is a desktop agent for knowledge work. It runs on your computer and works with your files under your permission, on the model you pick, with your own key.
Below: three ordinary jobs, given to it on 19 September 2026, shown exactly as they happened — the ask, the work, the file, and what it cost — and one of them run again four days later on another vendor's model.
Free for local use · Bring your own model keys · Nothing to upload
“Can we pay the machine deposit in November without the bank?”
Pieter owns a 27-person precision-machining company near Eindhoven. In a folder: twenty months of Rabobank export (in Dutch), the open customer invoices and supplier bills from the ERP, and a page of notes — a €485,000 machine on order, a €300,000 loan the bank has offered, a decision due by 31 October, and a floor of €150,000 he never wants to go under.
Hi. I own a small precision machining company near Eindhoven, 27 people … Cash has been tight since the spring and I want to see it coming instead of hearing it from the bank. In this folder you will find our business account export from the bank for the last 20 months (the CSV_A file, it is in Dutch), the open customer invoices and open supplier bills from our ERP, and notes.md … Please build me a cash flow forecast: week by week for the next 13 weeks and month by month for the next 12 months, in a spreadsheet I can update every month … Show me the bank balance over the next 13 weeks as a chart, with my 150,000 floor drawn on it. Then tell me in plain words when the cash is lowest, whether we can pay the machine deposit in November without the bank, and what I should do.
The ask, as typed · the full text is in the record
The work · GPT-5.6 Sol, thinking on High · one approval, then ten minutes on its own



The file · Cash_flow_forecast.xlsx, and the chart he asked for


| Same brief, same folder, two models | Clock | Cost | What it said | Against the fixture's truth |
|---|---|---|---|---|
| GPT-5.6 Luna · xHigh | 11 min 53 s | $0.19 | “The November deposit is affordable without the loan; the 13-week low is €381,561.” Recommended the loan anyway, for February. | Wrong on the question that mattered: the deposit week dips below zero without the overdraft, and the spring trough is about €250k deep, not €13k. |
| GPT-5.6 Sol · High | 9 min 52 s | $2.83 | “Without the loan, cash falls to €135,233 on 13 November — payable, but below your floor. Accept the loan before 31 October; you still need about €116k beyond it.” | Right fortnight, floor breach said plainly, the February hole found, the loan recommended in time. The deepest month lands two months late. |
What left the machine: your turns and what the model read from the four files, to api.openai.com — 38 requests (Luna), 26 (Sol). Nothing to anyone else. The bank export stayed on disk.
Recorded 19 Sep 2026 · run 2026-09-19-site-cases · Safata.app build of 16:19, the chart frame on the 19:29 rebuild · one key: OpenAI · the truth is the fixture's
“The board meets on the 22nd and I owe them the Q2 review.”
The managing director of a 40-person advisory firm. In the folder: 1,400 rows straight out of the CRM — re-entered deals, a renamed stage, UK clients billed in pounds — and his notes on what the board will ask. He wants ten slides with charts, a speaker note under each, and an appendix of what was wrong with the data and what was done about it.
Hi. I am the managing director of Northgate Advisory, a 40-person B2B advisory firm. Our board meets on the 22nd and I owe them the Q2 review. In this folder you will find pipeline_export.csv, the raw export from our CRM for the last three quarters, and notes.md where I explain what the board will ask about and what I know is wrong with the data. Please produce the Q2 review: what happened versus Q1 and Q4, by segment and by rep, the health of the pipeline going into Q3, and the three things the board should take away, as a ten-slide deck with charts and a speaker note under each slide, plus an appendix listing what you found wrong in the data and what you did about it. Read the notes first.
The ask, as typed
The work · GPT-5.6 Luna, thinking on xHigh · six minutes



| Model | Clock | Cost | What it said | Against the fixture's truth |
|---|---|---|---|---|
| GPT-5.6 Luna · xHigh | 6 min 7 s | $0.08 | Q2 €10.05m against a €10.5m target; Q1 €10.95m once the one-off Helix deal is set aside; €43.3m open going into Q3, 564 deals, €24.5m of it over 90 days old. | Pipeline exact; Q1-ex-Helix within €30k; Q2 3% under the truth (€10.36m) — it removed a few real deals with the re-entered ones, and says so in the appendix. |
What left the machine: your turns and what the model read from the export, to api.openai.com — 22 requests. Nothing to anyone else.
Recorded 19 Sep 2026 · run 2026-09-19-site-cases · the same brief ran on three models on 13–14 September
“I need a memo I can forward to my auditor.”
The CFO of a 50-person Dutch software company that invoices customers in Germany, France, Belgium and Poland. Mandatory e-invoicing is arriving in all four, on different dates and different rules, and he needs to know what his company has to do, by when, as a foreign supplier — with the tax authority, the gazette or an official FAQ behind every line.
Hi. I am the CFO of a 50-person software company based in the Netherlands. We invoice business customers in Germany, France, Belgium and Poland. I keep hearing about mandatory e-invoicing in these countries and I need a memo I can rely on. For each of the four countries: what becomes mandatory and when - issuing versus receiving, B2B versus B2G - which thresholds or company sizes it applies to, which formats and platforms are required, and what my company specifically has to do by which date as a foreign supplier invoicing into that country. Every claim must cite a primary source … with the date of the source. Where the rules are still uncertain or pending, say so rather than guessing. One memo, Word or HTML, that I can forward to my auditor.
The ask, as typed
The work · GPT-5.6 Luna, thinking on xHigh · four countries at once



| Model | Clock | Cost | What it said | What the run showed |
|---|---|---|---|---|
| GPT-5.6 Luna · xHigh | 17 min, 7 of them waiting on two approvals | $0.34 | The four regimes do not line up: Germany, Belgium and Poland key off local establishment; France adds a reporting track for foreign businesses liable for French VAT; public-sector invoicing is the exception everywhere. | One of the four workers lost its model stream for five minutes (the vendor's side); the harness aborted that step and said so, and the parent finished Poland itself. That line is on the page because it happened. |
What left the machine: your turns, to api.openai.com — 47 requests from the memo and its four workers, the searches inside them. Nothing to anyone else.
Recorded 19 Sep 2026 · run 2026-09-19-site-cases
“Someone sent me a Kimi K3 benchmark. Is it any good on our data?”
The same managing director, four days after the board deck. A colleague has forwarded him a benchmark chart for Kimi K3, Moonshot's model. He has the folder, the brief, and a deck he has already read line by line. He does not want the benchmark; he wants to know what this model does with his CRM export, and what it costs to find out.
He wrote no new brief. He pasted Moonshot's key into Settings, ticked Kimi K3, opened a new session on it, and pasted the brief of 19 September, word for word.
Hi. I am the managing director of Northgate Advisory, a 40-person B2B advisory firm. Our board meets on the 22nd and I owe them the Q2 review. In this folder you will find pipeline_export.csv, the raw export from our CRM for the last three quarters, and notes.md where I explain what the board will ask about and what I know is wrong with the data. Please produce the Q2 review … as a ten-slide deck with charts and a speaker note under each slide, plus an appendix listing what you found wrong in the data and what you did about it. Read the notes first.
The ask, as typed · the brief of case 2, verbatim, in a new session on Kimi K3
The switch · one key, pasted that afternoon · two clicks


The work · Kimi K3, thinking at its one level, max · two approvals, then 29 minutes on its own



The file · Northgate Q2 2026 board review.pptx, and the appendix it wrote alongside


| Same brief, same folder, two vendors | Clock | Cost | What it said | Against the fixture's truth |
|---|---|---|---|---|
| GPT-5.6 Luna · xHigh · 19 Sep | 6 min 7 s | $0.08 | Q2 €10.05m against a €10.5m target; Q1 €10.95m once the one-off Helix deal is set aside; €43.3m open going into Q3, 564 deals, €24.5m of it over 90 days old. | Pipeline exact; Q1-ex-Helix within €30k; Q2 3% under the truth (€10.36m) — it removed a few real deals with the re-entered ones, and says so in the appendix. |
| Kimi K3 · max · 23 Sep | 52 min, 23 of them waiting on two cards | $1.61 | “€10.27m won against the €10.5m target, 97.8% — a small miss; down 4.3% on Q1 like-for-like, up 51% on Q4; 145 wins, a record; €43.3m open, half of it older than 120 days.” Speaker notes in the notes pane; the two ×100 corrections flagged for ops with their sensitivity. | Closest of the four sittings on this brief: Q2 0.9% under the truth, Q4 3.9% under — the seven undated wins it excluded, and said so. All three ×100 rows scaled to the planted values; the twelve re-entries counted once; pipeline exact. Slower: 29 minutes of work against Luna's six. |
What left the machine: your turns and what the model read from the export, to api.moonshot.ai — 25 requests; and fourteen fetches of Python libraries from the Pyodide index, which the second card named and he approved. Nothing to anyone else. The CRM export stayed on disk. On 19 September the same brief was 22 requests to api.openai.com.
Kimi K3 in Safata today: thinking always on at maximum, so it is slow and priced like the big models; no vision, so it built its charts without looking at them; no search of its own, so a research job needs a search key in Settings. The receipt above is with all three true.
Recorded 23 Sep 2026 · run 2026-09-23-site-kimi · Safata.app build of 10:02 · one new key: Moonshot, added that afternoon · the truth is the fixture's · the same brief ran on three models on 13–14 September
What you just saw
- Any model, one folder each. Case 1 ran on two models of one vendor; case 2 ran again four days later on another vendor's, through a key pasted that afternoon. The benchmark that matters is last week's brief. Every major vendor, and hundreds more via OpenRouter, on your own keys.
- The file beside the conversation. The workbook, the deck and the memo opened on the canvas as they landed, computed and readable, while the agent was still talking. You keep the file, not the chat.
- One approval, the files named. Each job asked once, before anything ran, and listed exactly what it would read and write. Then it worked alone. See the plan before it runs.
- One line says what left. Under every session: the model's API and the pages it read. The bank export, the CRM export and the invoices stayed on disk. Everything stays here, except the turn you send.
It does other jobs too: a research memo with its sources · two hundred documents into one table · a CRM read through a read-only token · a support queue drafted from the manual · a report on a schedule, to a channel. How it works →
On your machine, it's free.
Safata as you see it here, local and on your own keys, is free for local use, at home and at work, employed or freelance. A session costs whatever your model vendor charges your own key: the three jobs above cost $4.11 together, most of it the big model's ten minutes and its two charts, and the rerun on Kimi K3 $1.61. I add no markup and resell nothing. A hosted plan and organizations are paid; nobody sells Safata itself.
Early build
Safata is a native desktop app: small, fast, built on Tauri. It opens on a keystroke and works against your disk at disk speed; open three windows and three agents run at once, on three different models if you like.
Status An early build, signed and notarized by Apple, so it opens like any other app. If something breaks, tell me; I read every one.
Version 0.1.3, for Macs with Apple silicon on macOS 13 or later. Windows comes later, and there is no Intel Mac build. It never checks for updates, because nothing in it phones home: watch the releases page and download the next one the same way. You'll need at least one API key from a model provider; getting one takes about two minutes.