safata Download
An agent for knowledge work · for the Mac

Give it a job.
Open the finished file.

Safata is a desktop agent for knowledge work. It runs on your computer and works with your files under your permission, on the model you pick, with your own key.

Below: three ordinary jobs, given to it on 19 September 2026, shown exactly as they happened — the ask, the work, the file, and what it cost — and one of them run again four days later on another vendor's model.

Download for Mac Read the three cases Apple silicon · macOS 13 or later

Free for local use · Bring your own model keys · Nothing to upload

Case 1 · a cash-flow forecast, run on two models

“Can we pay the machine deposit in November without the bank?”

Pieter owns a 27-person precision-machining company near Eindhoven. In a folder: twenty months of Rabobank export (in Dutch), the open customer invoices and supplier bills from the ERP, and a page of notes — a €485,000 machine on order, a €300,000 loan the bank has offered, a decision due by 31 October, and a floor of €150,000 he never wants to go under.

Hi. I own a small precision machining company near Eindhoven, 27 people … Cash has been tight since the spring and I want to see it coming instead of hearing it from the bank. In this folder you will find our business account export from the bank for the last 20 months (the CSV_A file, it is in Dutch), the open customer invoices and open supplier bills from our ERP, and notes.md … Please build me a cash flow forecast: week by week for the next 13 weeks and month by month for the next 12 months, in a spreadsheet I can update every month … Show me the bank balance over the next 13 weeks as a chart, with my 150,000 floor drawn on it. Then tell me in plain words when the cash is lowest, whether we can pay the machine deposit in November without the bank, and what I should do.

The ask, as typed · the full text is in the record

The work · GPT-5.6 Sol, thinking on High · one approval, then ten minutes on its own

The plan card before anything runs: it reads the bank export, the invoices, the bills and the notes, writes Cash_flow_forecast.xlsx, and asks once — Approve.
It asks once. Before anything runs, one card names every file it may read and the one it may write. Approve, and nothing below asks again.
The ladder of steps in plain English: identifying recurring payroll, tax, rent, lease and loan cash items; stress-testing Brainport's 75-day collections.
You can read what it is doing. Each step is a sentence — Identifying recurring payroll, tax, rent, lease and loan cash items — not a command line.
The finished reply beside the workbook opened on the canvas: 13 sheets, 34,367 formulas computed here, the decision dashboard on top.
The file opens beside the conversation. 13 sheets, 34,367 live formulas, computed here — the decision dashboard on top.

The file · Cash_flow_forecast.xlsx, and the chart he asked for

The 13-week balance chart on the canvas: the no-loan and with-loan lines against the dashed €150,000 floor; the no-loan line dips below it in mid-November.
Without the loan, cash falls below €150k in November. The chart is drawn on the canvas from the workbook's own numbers; the dashed line is his floor.
The workbook's Summary sheet on the canvas: lowest date 13 Nov 2026, lowest balance €135,233, deposit decision: payable but below the cash floor.
The answer, in the sheet and in words. Lowest point 13 November at €135,233 — the deposit is payable, but not while keeping the floor. Take the loan before 31 October.
Same brief, same folder, two modelsClockCostWhat it saidAgainst the fixture's truth
GPT-5.6 Luna · xHigh11 min 53 s$0.19“The November deposit is affordable without the loan; the 13-week low is €381,561.” Recommended the loan anyway, for February.Wrong on the question that mattered: the deposit week dips below zero without the overdraft, and the spring trough is about €250k deep, not €13k.
GPT-5.6 Sol · High9 min 52 s$2.83“Without the loan, cash falls to €135,233 on 13 November — payable, but below your floor. Accept the loan before 31 October; you still need about €116k beyond it.”Right fortnight, floor breach said plainly, the February hole found, the loan recommended in time. The deepest month lands two months late.

What left the machine: your turns and what the model read from the four files, to api.openai.com — 38 requests (Luna), 26 (Sol). Nothing to anyone else. The bank export stayed on disk.

Case 2 · a board deck from a dirty CRM export

“The board meets on the 22nd and I owe them the Q2 review.”

The managing director of a 40-person advisory firm. In the folder: 1,400 rows straight out of the CRM — re-entered deals, a renamed stage, UK clients billed in pounds — and his notes on what the board will ask. He wants ten slides with charts, a speaker note under each, and an appendix of what was wrong with the data and what was done about it.

Hi. I am the managing director of Northgate Advisory, a 40-person B2B advisory firm. Our board meets on the 22nd and I owe them the Q2 review. In this folder you will find pipeline_export.csv, the raw export from our CRM for the last three quarters, and notes.md where I explain what the board will ask about and what I know is wrong with the data. Please produce the Q2 review: what happened versus Q1 and Q4, by segment and by rep, the health of the pipeline going into Q3, and the three things the board should take away, as a ten-slide deck with charts and a speaker note under each slide, plus an appendix listing what you found wrong in the data and what you did about it. Read the notes first.

The ask, as typed

The work · GPT-5.6 Luna, thinking on xHigh · six minutes

The plan card for the board review: it reads notes.md and pipeline_export.csv and writes the deck — one approval.
One approval. Two files it may read, one it may write.
The save-the-deck card listing every slide title, and the steps above it: checking owner-change re-entries, listing impossible negative deal values for the appendix.
It read the notes. Check for owner-change re-entries that are separate CRM records · List impossible negative deal values for the data-quality appendix. The deck card names every slide before it saves.
The finished reply beside the deck on the canvas: the title slide, then a slide reading Q2 missed target by €0.45m despite more wins, with its bar chart and speaker note.
The deck, slide by slide, beside the conversation. Ten slides plus two appendix slides, each with its speaker note.
ModelClockCostWhat it saidAgainst the fixture's truth
GPT-5.6 Luna · xHigh6 min 7 s$0.08Q2 €10.05m against a €10.5m target; Q1 €10.95m once the one-off Helix deal is set aside; €43.3m open going into Q3, 564 deals, €24.5m of it over 90 days old.Pipeline exact; Q1-ex-Helix within €30k; Q2 3% under the truth (€10.36m) — it removed a few real deals with the re-entered ones, and says so in the appendix.

What left the machine: your turns and what the model read from the export, to api.openai.com — 22 requests. Nothing to anyone else.

Case 3 · a memo with a dated source on every claim

“I need a memo I can forward to my auditor.”

The CFO of a 50-person Dutch software company that invoices customers in Germany, France, Belgium and Poland. Mandatory e-invoicing is arriving in all four, on different dates and different rules, and he needs to know what his company has to do, by when, as a foreign supplier — with the tax authority, the gazette or an official FAQ behind every line.

Hi. I am the CFO of a 50-person software company based in the Netherlands. We invoice business customers in Germany, France, Belgium and Poland. I keep hearing about mandatory e-invoicing in these countries and I need a memo I can rely on. For each of the four countries: what becomes mandatory and when - issuing versus receiving, B2B versus B2G - which thresholds or company sizes it applies to, which formats and platforms are required, and what my company specifically has to do by which date as a foreign supplier invoicing into that country. Every claim must cite a primary source … with the date of the source. Where the rules are still uncertain or pending, say so rather than guessing. One memo, Word or HTML, that I can forward to my auditor.

The ask, as typed

The work · GPT-5.6 Luna, thinking on xHigh · four countries at once

Four sub-tasks running side by side — Researching Germany, France, Belgium, Poland — each with what it is fetching, its tool count and its elapsed time.
Four countries, four workers. One per country, each with its own brief; the conversation shows what each is doing and how long it has taken.
Search rows naming the sites read: site:ksef.podatki.gov.pl and site:gesetze-im-internet.de — the tax authorities themselves.
You can see where it looked. site:ksef.podatki.gov.pl, site:gesetze-im-internet.de — the tax authorities themselves, named in the row.
The finished reply: the memo's scope, its 24-entry primary-source register, and the file receipt.
One HTML file he can forward, with a 24-entry register of dated primary sources, checked link by link before hand-over.
ModelClockCostWhat it saidWhat the run showed
GPT-5.6 Luna · xHigh17 min, 7 of them waiting on two approvals$0.34The four regimes do not line up: Germany, Belgium and Poland key off local establishment; France adds a reporting track for foreign businesses liable for French VAT; public-sector invoicing is the exception everywhere.One of the four workers lost its model stream for five minutes (the vendor's side); the harness aborted that step and said so, and the parent finished Poland itself. That line is on the page because it happened.

What left the machine: your turns, to api.openai.com — 47 requests from the memo and its four workers, the searches inside them. Nothing to anyone else.

Case 2, again · the same deck, four days later, on a model from another vendor

“Someone sent me a Kimi K3 benchmark. Is it any good on our data?”

The same managing director, four days after the board deck. A colleague has forwarded him a benchmark chart for Kimi K3, Moonshot's model. He has the folder, the brief, and a deck he has already read line by line. He does not want the benchmark; he wants to know what this model does with his CRM export, and what it costs to find out.

He wrote no new brief. He pasted Moonshot's key into Settings, ticked Kimi K3, opened a new session on it, and pasted the brief of 19 September, word for word.

Hi. I am the managing director of Northgate Advisory, a 40-person B2B advisory firm. Our board meets on the 22nd and I owe them the Q2 review. In this folder you will find pipeline_export.csv, the raw export from our CRM for the last three quarters, and notes.md where I explain what the board will ask about and what I know is wrong with the data. Please produce the Q2 review … as a ten-slide deck with charts and a speaker note under each slide, plus an appendix listing what you found wrong in the data and what you did about it. Read the notes first.

The ask, as typed · the brief of case 2, verbatim, in a new session on Kimi K3

The switch · one key, pasted that afternoon · two clicks

Settings › Model: the catalog with OpenAI's seven rows and, under Moonshot, Kimi K3 ticked — 1M context, $3 in and $15 out — and the line: tick a model to put it on your menu.
One key, one row. Paste the vendor's key in Settings and its models appear in the catalog, priced, beside the ones you already had. Tick one and it is on your menu.
The session chip's picker: GPT-6 Astra, GPT-6 Sol, GPT-6 Luna, GPT-5.6 Luna marked default, Kimi K3, and All models…
The menu, on the chip. New sessions still open on the model he pinned. This one opens on Kimi K3, and the sessions he already did stay on the model they were written in.

The work · Kimi K3, thinking at its one level, max · two approvals, then 29 minutes on its own

The ladder of steps: checking date formats, amount anomalies and duplicate deal IDs; hunting the 9.7M outlier, negative amounts and re-entered duplicate deals; confirming the re-entered deals via the -R suffix; one crashed step and its retry; building the cleaned dataset; profiling the Q3 pipeline.
The same job, read differently. Hunting the 9.7M outlier, negative amounts, and re-entered duplicate deals · Confirming the re-entered deals via the -R suffix. One step crashed and was retried; it is on the frame because it happened.
The autopilot card: build the Northgate Q2 board review deck with speaker notes, plus the data-issues appendix, and verify both — read and grep the two inputs, write and edit the deck and the appendix, spend up to $4.00 — Approve, Deny.
It asks once, after fourteen minutes of reading: the two files it may read, the two it may write, the spend it may reach. Approve, and the build runs alone.
The floor: Safata wants to use its python tool — it fetches matplotlib and python-pptx from the Pyodide index at cdn.jsdelivr.net into Safata's own Python, nothing is installed into this computer's own — Once, This session, Always, This task, Deny.
The floor, under the approval. Building a deck needs two libraries Safata's own Python fetches from the web, and a fetch leaves the machine, so it asked again and named the host. Nothing is installed on the computer.

The file · Northgate Q2 2026 board review.pptx, and the appendix it wrote alongside

The finished reply's On the data paragraph and the two file cards, beside the deck on the canvas: the title slide and slide 2, €10.27m won — 98% of target; the miss is smaller than the Q1 comparison suggests.
Twelve slides beside the conversation, the ten-slide review and a two-slide data appendix — and the speaker notes in the notes pane, which the 19 September deck printed on the slide face. The two unit corrections are marked in the deck as pending ops confirmation.
Slides 3 and 4 on the canvas: where we won, by segment; and Anna carried the quarter, Elena is second, Ben's fall has a known cause — a bar chart by seller with Q1 and Q2.
The board's two named questions, answered by name. Ben's −57% traced to the framework bid he was moved to in May; Elena's three straight quarters, with her Q1 netted for the deals reassigned to her.
Same brief, same folder, two vendorsClockCostWhat it saidAgainst the fixture's truth
GPT-5.6 Luna · xHigh · 19 Sep6 min 7 s$0.08Q2 €10.05m against a €10.5m target; Q1 €10.95m once the one-off Helix deal is set aside; €43.3m open going into Q3, 564 deals, €24.5m of it over 90 days old.Pipeline exact; Q1-ex-Helix within €30k; Q2 3% under the truth (€10.36m) — it removed a few real deals with the re-entered ones, and says so in the appendix.
Kimi K3 · max · 23 Sep52 min, 23 of them waiting on two cards$1.61“€10.27m won against the €10.5m target, 97.8% — a small miss; down 4.3% on Q1 like-for-like, up 51% on Q4; 145 wins, a record; €43.3m open, half of it older than 120 days.” Speaker notes in the notes pane; the two ×100 corrections flagged for ops with their sensitivity.Closest of the four sittings on this brief: Q2 0.9% under the truth, Q4 3.9% under — the seven undated wins it excluded, and said so. All three ×100 rows scaled to the planted values; the twelve re-entries counted once; pipeline exact. Slower: 29 minutes of work against Luna's six.

What left the machine: your turns and what the model read from the export, to api.moonshot.ai — 25 requests; and fourteen fetches of Python libraries from the Pyodide index, which the second card named and he approved. Nothing to anyone else. The CRM export stayed on disk. On 19 September the same brief was 22 requests to api.openai.com.

Kimi K3 in Safata today: thinking always on at maximum, so it is slow and priced like the big models; no vision, so it built its charts without looking at them; no search of its own, so a research job needs a search key in Settings. The receipt above is with all three true.

What you just saw

  1. Any model, one folder each. Case 1 ran on two models of one vendor; case 2 ran again four days later on another vendor's, through a key pasted that afternoon. The benchmark that matters is last week's brief. Every major vendor, and hundreds more via OpenRouter, on your own keys.
  2. The file beside the conversation. The workbook, the deck and the memo opened on the canvas as they landed, computed and readable, while the agent was still talking. You keep the file, not the chat.
  3. One approval, the files named. Each job asked once, before anything ran, and listed exactly what it would read and write. Then it worked alone. See the plan before it runs.
  4. One line says what left. Under every session: the model's API and the pages it read. The bank export, the CRM export and the invoices stayed on disk. Everything stays here, except the turn you send.

It does other jobs too: a research memo with its sources · two hundred documents into one table · a CRM read through a read-only token · a support queue drafted from the manual · a report on a schedule, to a channel. How it works →

The deal

On your machine, it's free.

Safata as you see it here, local and on your own keys, is free for local use, at home and at work, employed or freelance. A session costs whatever your model vendor charges your own key: the three jobs above cost $4.11 together, most of it the big model's ten minutes and its two charts, and the rerun on Kimi K3 $1.61. I add no markup and resell nothing. A hosted plan and organizations are paid; nobody sells Safata itself.

Early build

Safata is a native desktop app: small, fast, built on Tauri. It opens on a keystroke and works against your disk at disk speed; open three windows and three agents run at once, on three different models if you like.

Status An early build, signed and notarized by Apple, so it opens like any other app. If something breaks, tell me; I read every one.

Version 0.1.3, for Macs with Apple silicon on macOS 13 or later. Windows comes later, and there is no Intel Mac build. It never checks for updates, because nothing in it phones home: watch the releases page and download the next one the same way. You'll need at least one API key from a model provider; getting one takes about two minutes.