The interesting part of a screen-aware assistant is not what it can see. It is what it declines to see, and what it refuses to claim.
Screen reading should never be something you discover is already running. Screen Context is off by default, and the Screen Agent sits on top of it — with Screen Context off, the agent cannot run at all. Settings says so outright instead of failing quietly.
Reads the text of the active window. Off on a fresh install, and everything on this page depends on it.
The part that forms an opinion and occasionally says something. Requires the switch above; the settings panel refuses rather than pretending.
A single global on/off is not enough. You pick a mode: a blocklist, where everything is readable except what you exclude, or an allowlist, where only the apps you name are readable and nothing else is.
Password managers are merged into the blocklist automatically, and the blocklist beats the allowlist. Allowlist your password manager — deliberately or by accident — and it still cannot be read. The safe outcome does not depend on you configuring it correctly.
Reading recognised text is one thing. Sending a picture of your window to a vision model is a different thing. Bundling those into one toggle is the mistake this avoids — Visual mode is its own consent, never rolled into another.
Text only. It will not claim things only eyes can verify — a disabled button, a highlighted field. Nothing leaves your Mac as an image.
One downscaled image of the allowed, focused window may go to the cloud vision model — only when text alone cannot answer. Never stored, never logged.
A model handed only text will happily assert that a button is greyed out. It cannot know that. So every visible claim is bound to evidence the agent actually read, and each card cites where it read it. Colour is reserved for something genuinely being wrong, not used to make ordinary remarks look urgent.
As of 1.3.27 the agent looks at your open tasks and confirmed facts before it looks at your screen, instead of treating every glance as a fresh universe. It had the ability to look those up for a while and simply never used it.
The second way these assistants fail is not privacy — it is noise. One that comments on everything gets switched off within a day.
Quiet, Balanced or Frequent, plus a cooldown. You set how often it is allowed to speak at all.
Stops the interruptions without turning the feature off. Held comments keep collecting while you are heads-down.
Cards on screen are short-lived on purpose. Every comment lands in the inbox anyway — including the ones the agent decided to hold back rather than show you — and the menu bar carries an unread count, so a missed card is not a lost one. Press Return on any card to carry it into the conversation with its source context intact, or mark it not helpful and say what was wrong with it.
No. Screen Context is off by default and the Screen Agent depends on it — with Screen Context off the agent cannot run. Visual mode is a third, separate switch, also off.
Yes. Blocklist mode reads everything except the apps you exclude; allowlist mode reads only the apps you name. Password managers are force-excluded and the blocklist beats the allowlist, so allowlisting one still cannot expose it.
Only with Visual mode on, only one downscaled image of the allowed focused window, and only when recognised text alone cannot answer. Never stored, never logged. With Visual mode off, no image leaves your Mac.
In text-only mode it will not claim anything only eyes could verify — that a button is disabled, that a field is highlighted. Every card cites the evidence it actually read, and if it cannot point at evidence it stays quiet.
No. Cards are deliberately short-lived on screen, but every comment lands in the inbox — including the ones the agent held back — with an unread count in the menu bar. ⌘⌥O opens it.
It needs access to a language model: a local model, your own API key, or Pro. The screen reading itself happens on your Mac.
Download MetaWhisp free. Everything on this page is off until you switch it on, scoped to the apps you name, and it will not tell you anything it cannot show you the source for.
Download MetaWhispmacOS 14+ · Apple Silicon · screen reading off by default
MetaWhisp is a free, open-source, on-device voice-to-text (dictation) app for macOS. It uses Whisper large-v3-turbo running on Apple Neural Engine. Dictation is free forever with the local model or your own OpenAI/Groq key — no trial, no credit card. Optional Pro ($7.77/mo or $30/yr) adds cloud processing with zero setup.
Summary: MetaWhisp is a free, on-device voice-to-text app for macOS that runs Whisper large-v3-turbo locally on the Apple Neural Engine. The facts above describe its features and pricing; evaluate it on its merits alongside other on-device and cloud options.