ozen
ozen (אוזן, “ear”) is a local meeting copilot for macOS. While you’re in a Zoom, Meet or Teams call, ozen listens, transcribes the call and the room in Hebrew and English, knows who is speaking, and lets an AI assistant (Claude Code or Hermes) read the meeting and see your screen so it can answer questions as the meeting happens.
Everything runs on your Mac. No bot joins the meeting and no audio leaves the machine.
What you get
- A live transcript of the call and of people in the room, one line per utterance, with a speaker name on every line. Lines appear about 15–30 seconds after they’re spoken.
- Speaker names that stick. ozen recognizes voices it has heard before. Tell it once who said a line and it relabels the rest of the meeting, and remembers that person next time.
- A transcript that learns. Fix a misheard word and ozen learns the word; fix the same mistake twice and it corrects it by itself from then on.
- Echo and noise filtered out. Remote voices leaking from your speakers into the mic, the computer reading text aloud, and a video playing nearby don’t pollute the transcript.
- Overlapping speech separated. When two people talk at once, each gets their own line.
- Recording on your terms. Record always, only during meetings, or by place (auto on at the office, off at home).
- AI on tap. Hand one or more past meetings, or the one happening right now, to an AI agent with one click. Any MCP-capable agent can also read and edit ozen’s data directly.
A transcript line looks like this:
[15:06:52] Ofek Gabay (room): אימבדין שלי, אני לא מוצא את זה
[15:09:41] S2 (call): Let's move to the launch timeline.
(call) lines came from the meeting app, (room) lines from your microphone. S2 is a voice ozen hasn’t been
told a name for yet.
Where to start
- Install ozen and grant its permissions.
- Record your first meeting.
- Learn the menu bar panel, where you’ll spend most of your time.
Before recording other people, read Privacy.
Install
Requirements
- A Mac with Apple Silicon running macOS 15 or later
- Xcode command line tools (
xcode-select --install), forswiftc - Rust (
cargo) - uv (Python runner)
ffmpeg(brew install ffmpeg)- Read access to the ozen repository and to your voices registry
Build and install
git clone https://github.com/ozenhq/ozen ~/ozen
git clone https://github.com/tupe12334/voices-embedding-registry ~/ozen/voices
cd ~/ozen && cargo run --release -- app
The last command builds the ozen command-line tool (~/ozen/target/release/ozen) and installs
Ozen.app into ~/Applications.
Open Ozen from Spotlight, Launchpad or Finder like any app. It lives in the menu bar as an ear icon and has no Dock icon.
Tip: put the CLI on your
PATHso you can typeozeninstead of~/ozen/target/release/ozen: addalias ozen=~/ozen/target/release/ozento~/.zshrc. The rest of these docs writeozen.
Permissions
The first time you press Start, macOS asks Ozen for two permissions. Grant both in System Settings › Privacy & Security, then press Start again:
| Permission | Why |
|---|---|
| Screen & System Audio Recording | To capture audio from meeting apps and other apps (ScreenCaptureKit). ozen does not record video. |
| Microphone | To hear people in the room, including you. |
| Location Services (optional) | Asked only when you first locate a place. Without it, places never match and recording follows Always / Meetings. |
The build signs the app with a local self-signed certificate (kept in
~/Library/Keychains/ozen-signing.keychain-db), so permissions survive rebuilds.
First-run downloads
The speech models download the first time they’re needed: Whisper (Hebrew and English) and the voice-recognition model on the first transcription, and the ~640 MB speech separator in the background after the first start. Until the separator is ready, two people talking at once stay merged in one line. While models download the panel may say the transcriber is catching up; that’s expected.
Updating
cd ~/ozen && git pull && cargo run --release -- app
Rebuilding never stops a recording in progress.
Another Mac
Run the same three install commands on the other Mac. Voices you tag on either Mac reach the other within a few minutes; see Using several Macs.
Your first meeting
- Start recording. Click the ear icon in the menu bar and press Start. The icon turns into a red filled ear while ozen records.
- Join your call in Zoom, Meet (in Chrome), Teams, Slack, FaceTime or Discord. Nothing to configure: ozen hears meeting apps directly, not through your speakers.
- Watch the transcript. Open the panel from the ear icon. New lines arrive every few seconds, about 15–30 seconds behind speech. Hebrew lines read right to left.
- Name the speakers. Voices ozen doesn’t know show as
S1,S2, … Click a speaker name and pick or type who it is. ozen relabels that voice everywhere in the meeting and recognizes it next time. See Who is speaking. - Fix a word if one was misheard: click the line’s text and type the correction. See Fixing the transcript.
- Ask AI. Press Ask AI to open Claude Code or Hermes on the meeting happening now. Ask “what did we decide about the launch date?” or “what’s on my screen?”. See Working with AI agents.
- Stop. Press Stop. ozen finishes transcribing what it has already heard (the icon shows an hourglass), then stops.
Don’t want to press Start every time?
Switch the panel to Meetings mode and ozen starts by itself whenever a meeting app uses the microphone, and stops shortly after the call ends. See When ozen records.
The menu bar panel
Ozen lives in the menu bar as an ear icon.
- Click the icon for the panel.
- Right-click for a menu with the same controls (Start / Pause / Resume / Stop, Always / Meetings, Places…).
The icon
| Icon | Meaning |
|---|---|
| Red filled ear | Recording |
| Monochrome ear | Stopped |
| Pause symbol | Paused |
| Hourglass | Stopped, finishing transcription of audio already heard |
When not recording, the panel says Not recording and why (paused, finishing transcription, or a problem).
Controls
| Button | What it does |
|---|---|
| Start | Start recording the call, the room microphone and other apps’ audio. |
| Pause / Resume | Stop and restart listening. The transcriber stays loaded, so resuming is instant. |
| Stop | Stop listening, finish transcribing queued audio (up to 2 minutes), then unload everything. |
| Always / Meetings | Record mode; see When ozen records. |
| Places… | Place-based rules; see Places. |
| Review N | Jump to the line ozen is least sure who said; see Review. |
| Ask AI | Open an AI agent on the meeting happening now; see Working with AI agents. |
Problems show as warnings at the top of the panel: recording blocked by a missing permission, a silent microphone, a screen that’s asleep, or a transcriber that’s behind. See Troubleshooting.
Views
Transcript
The live transcript, refreshed every 2 seconds. Each line shows its time, speaker and text.
- Click a speaker name to say who really said it, or to ignore that voice.
- Click the text to correct it.
- An orange ? marks a line ozen is unsure about.
- Grey lines are ignored voices.
Timeline
One lane per speaker, with their total talk time, and a bar for every line they spoke along a scrollable time axis. Overlapping speech overlaps here too.
- − / + zoom.
- Silences over 2 minutes shrink to a short break marker.
- Hover a bar to read the line; click it to jump to that line in the transcript.
Meetings
Every past meeting, newest first. A silence of 10 minutes or more starts a new meeting.
Select one or more (⌘-click or ⇧-click) and press Open to hand them to an AI agent; see Working with AI agents.
The footer shows speaker-recognition accuracy, now and when you started.
When ozen records
What ozen hears
While recording, ozen captures three audio streams:
| Stream | Source | Transcribed? |
|---|---|---|
| call | Meeting apps only: Zoom, Chrome, Teams, Slack, FaceTime, Discord | Yes, as (call) lines |
| mic | Your microphone, meaning everyone in the room | Yes, as (room) lines |
| local | Every other app, like a video or the say command | No. Used only to recognize when the computer itself is talking, so it isn’t mistaken for a person in the room |
Record modes
Pick a mode with the Always / Meetings toggle in the panel or the right-click menu.
Always. Records from Start until you pause or stop, including everything the room mic hears.
Meetings. Starts by itself when a meeting app (Zoom, Chrome/Meet, Teams, Slack, FaceTime, Discord) starts using the microphone, and stops 20 seconds after it releases it. This works for any call in those apps, with no calendar or plugin needed.
In either mode, pausing or starting by hand wins until the next meeting starts or ends. Relaunching Ozen never stops a recording in progress.
Places
Places switch recording by where your Mac is. Open Places… from the panel or the right-click menu.
Each place has a label, a location, a radius, and one of three settings:
| Setting | While you’re there |
|---|---|
| Auto record | Record (like Always) |
| Record meetings only | Record during calls (like Meetings) |
| Auto off | Don’t record |
While you’re inside a place’s radius, its setting replaces Always / Meetings. Arriving or leaving applies right away; a manual pause or start holds until then. Outside every place, your Always / Meetings choice applies.
Home and Work are there from the start with no location, so they do nothing until you set them. To set a place’s location:
- type its latitude and longitude,
- press Use current location,
- press Pick on map and click, or
- drag its pin on the map.
The first time, macOS asks for Location Services. Right after Ozen launches it waits for your location before any place decides anything.
The map uses OpenStreetMap; viewing it downloads map tiles for that area. Your location itself is never sent
anywhere. Places are saved in ~/ozen/places.json; see Files and data.
Who is speaking
ozen gives every line a speaker by its voice. Each utterance gets a voiceprint, which is compared with voices already heard, so one person keeps one label for the whole meeting.
- Known people show by name. They come from your voices registry, which grows as you tag.
- Unknown voices show as
S1,S2, … until you name them.
Naming a speaker (tagging)
Click any speaker name in the transcript and pick or type who really said that line. Then ozen:
- rebuilds that person’s voiceprint from every line tagged as them, in this meeting and earlier ones,
- relabels every untagged line with the new voiceprints, past and future,
- saves the voice to the registry, so your other Macs learn it too.
This takes a few seconds. A few tags per person early in a meeting fix most labels. To undo a tag, click the speaker name and choose Clear tag.
From the command line: ozen tag <line-id> "Dana Levi" (an empty name clears the tag).
Review
ozen scores every untagged line and queues the ones it’s least sure about: near the same-voice cutoff, or nearly tied between two people. They show an orange ? in the transcript.
Press Review N to jump to the most uncertain line and answer who said it. Review asks only about the last 10 minutes, because after that nobody remembers who said what.
Answering Review is the fastest way to improve recognition. In a test with four similar voices, 12 tags chosen by Review reached 100% accuracy on untagged lines, against 98% for 12 random tags.
Every tag also recalibrates how similar two clips must be to count as the same voice, and logs accuracy; the panel footer shows it next to the starting value.
Ignoring a voice
A video playing next to your Mac, or a radio in the room, isn’t part of the meeting. Click its speaker name and choose:
- Ignore this voice to ignore that line’s voice, or
- Ignore all N lines by S3 to mark every nearby line of that speaker at once.
Those lines, and earlier ones that sound like them, turn grey and leave the timeline, and ozen stops writing that voice from then on. Ignored voices stay on this Mac; they’re never shared to the registry.
To undo, tag the line as a person, or choose Clear tag.
From the command line: ozen ignore <line-id>....
People talking at once
When two people talk at once, in the room, on the call, or one on each, ozen separates the audio into one track per voice and writes each as its own line with its own speaker and time. In the timeline, those lines overlap.
Only utterances where voices disagree are separated, so a single speaker costs nothing extra. See Limits for what isn’t separated.
Echo
A mic utterance that mostly overlaps call or computer audio, in the same voice, is your speakers bleeding into the mic, not a person in the room, so ozen drops it. That covers the computer reading text aloud and remote voices leaking into the mic.
Someone in the room talking over the call keeps their line, because their voice doesn’t match what was playing.
Fixing the transcript
Correcting a line
Click a line’s text in the panel and type what was really said. Leave it empty to restore what ozen heard.
From the command line: ozen fix <line-id> "right text" (empty clears).
What ozen learns from fixes
- New words. Words your fixes add, like names, product terms and jargon, are hinted to the transcriber on every new line, most used first.
- Repeated corrections. Correct the same phrase the same way twice (say “פי אר” → “PR”) and ozen applies that correction to new lines automatically.
- A corrected line keeps what the transcriber originally heard, so if an automatic correction is wrong, fixing it back cancels it. A correction is also skipped whenever one of your fixes kept that phrase as right.
Each fix also keeps its audio with the right text, as a dataset for fine-tuning a model later. It stays on your Mac.
Vocabulary
~/ozen/vocab.txt lists names and terms the transcriber should spell right, like Kev, PR, code review. Edit
it freely, one term per line or comma-separated; it’s read on every line. Known people’s names are hinted automatically.
Agents can manage it too, with the list_vocab, add_vocab and remove_vocab MCP tools.
Languages
Each utterance is transcribed as either Hebrew or English, picked per utterance. Hebrew uses a Hebrew-trained Whisper model (ivrit.ai), English uses stock Whisper large-v3-turbo. English terms spoken inside Hebrew are the weakest spot; add them to the vocabulary.
Filler that Whisper tends to invent on silence (“Thank you.”, “תודה רבה”) is dropped.
Using several Macs
Voices are shared across your Macs through the voices registry, a private git repository that every Mac clones
into ~/ozen/voices.
- Start pulls the registry.
- While running, ozen pulls again every few minutes.
- Every tag rebuilds voiceprints from the newest registry and pushes when done.
So a person you tag on one Mac is recognized on the other within a few minutes.
Tagging on two Macs at once is safe. A push that loses the race rebuilds on top of the other Mac’s and pushes again, so both sets of tags survive.
Offline tags are committed locally and pushed with the next tag once you’re back online.
What stays on each Mac
Each Mac’s own tags and transcript never leave it; the registry holds only names and voiceprints. Ignored voices, places, fixes and recordings are local too. See Files and data.
Don’t edit
~/ozen/voicesby hand. Every retrain resets it to the remote before rebuilding, so hand edits are lost. Change voices by tagging.
Setting up another Mac
Follow Install on the new Mac, with the same registry URL.
Working with AI agents
ozen hands meetings to an AI coding agent, Claude Code or Hermes, in a folder
holding the transcripts plus an AGENTS.md (and CLAUDE.md) that tells the agent what’s there. The agent must be
installed and on your PATH.
The meeting happening now
Press Ask AI in the panel. ozen writes the current meeting (one whose last line is under 10 minutes old) to
~/ozen/context/live/ and starts the agent there in a new Terminal window.
While the meeting goes on, a background process rewrites that folder with the latest lines every 15 seconds, and
exits when the meeting ends. The agent is told to reread it, and that ozen look shows your screen, so you can
ask things like:
- “Summarize the last five minutes.”
- “What did Dana ask me to do?”
- “What’s on my screen right now, and does it match what they’re describing?”
From the command line: ozen live --open claude (or hermes); without --open it only writes the folder and
prints its path.
Past meetings
Switch the panel to Meetings, select one or more (⌘-click or ⇧-click), and press Open. ozen writes their
transcripts into a fresh folder under ~/ozen/context/ and starts the agent you choose there. The folder’s
claude.command or hermes.command reopens it later with a double-click.
Auto add with Kev also adds every other meeting that a local Kev server
(localhost:8009) judges part of the same project or topic. Its scores show before you pick the agent. Without
Kev running, this option does nothing.
From the command line:
ozen meetings # list meetings with their ids
ozen gather [--kev] ID... # write them to context/<now>/ and print the folder
ozen open DIR claude # or hermes, or finder
Screen and recent lines
ozen look [N] takes a screenshot of the main display (saved as ~/ozen/screen-small.png) and prints the last N
transcript lines (default 40), with tagged speakers and fixed text. Agents use it to answer “what’s on screen” or
“what was just said”. The app running it (like your terminal) needs Screen Recording permission.
Any MCP agent
ozen mcp gives any MCP client full access to meetings, transcript lines, people, places, vocabulary and
recording control. See MCP server.
Command line
The ozen binary is built to ~/ozen/target/release/ozen. It works from any directory. Most commands do what a
panel button does, so you can script ozen or drive it from an agent.
Recording
| Command | What it does |
|---|---|
ozen start | Start recording (same as Start). |
ozen pause | Pause; the transcriber stays loaded. |
ozen resume | Resume after a pause. |
ozen stop | Stop, finish transcribing queued audio (up to 2 minutes), then unload. |
ozen status | Print recording, paused, stopping or stopped. |
ozen health | Print one line per problem the panel warns about; nothing when all is well. See Troubleshooting. |
Transcript
| Command | What it does |
|---|---|
ozen show [N] | Last N lines, with tagged speakers and fixed text. |
ozen look [N] | Screenshot to ~/ozen/screen-small.png, then the last N lines (default 40). |
ozen fix <line-id> "right text" | Correct a line’s text; empty restores it. Relearns hint words and corrections. |
ozen tag <line-id> "Name" | Say who said a line; empty clears. Retrains and pushes the registry. |
ozen ignore <line-id>... | Mark lines as a voice to ignore. Retrains. |
ozen retrain | Rebuild voiceprints, labels, ignored voices and accuracy from all tags. Tagging does this for you. |
Meetings and agents
| Command | What it does |
|---|---|
ozen meetings | Past meetings, newest first: id, start, minutes, lines, first words. |
ozen gather [--kev] ID... | Write those meetings to ~/ozen/context/<now>/ and print the folder. --kev adds meetings Kev judges related. |
ozen live [--open claude|hermes] | Write the current meeting to ~/ozen/context/live/, keep it updated every 15 s until the meeting ends, and print the folder. --open starts that agent there. |
ozen open DIR claude|hermes|finder | Open a folder written by gather or live in that agent, or in Finder. |
ozen mcp | Run the MCP server on stdio. |
App
| Command | What it does |
|---|---|
ozen app | Build and install ~/Applications/Ozen.app. |
ozen bar | Build if needed, then open Ozen.app. |
Evaluation
| Command | What it does |
|---|---|
ozen eval [--vocab 0,10,30] [--repeat 0,1,2] [--real] [--fresh] | Score learning settings; see Tuning how fixes teach. |
uv run eval.py [N] | Transcribe your last N real chunks with stock, Hebrew and Hebrew+vocabulary models, to compare on your own speech. Run from ~/ozen. |
MCP server
ozen mcp is a Model Context Protocol server on stdio. It lets any MCP-capable
agent read and edit what ozen keeps: meetings, transcript lines, people, places, vocabulary, and recording.
Connect it
Claude Code:
claude mcp add -s user ozen -- ~/ozen/target/release/ozen mcp
Any other MCP client: command ~/ozen/target/release/ozen, argument mcp.
Concepts
- Lines are transcript lines: id, time, speaker, text. A line’s speaker is either a tag (set by you or an
agent), which wins, or the voiceprint’s guess.
textis the fixed text if the line was fixed;heardis what the transcriber originally wrote. - Meetings are runs of lines with under 10 minutes of silence between them. A meeting’s id is its start time in unix seconds.
- Notes are lines an agent creates, like a summary or an action item. They have no voice, so editing one just rewrites it.
- The speaker name Ignored marks a voice to drop.
Tools
Recording
| Tool | Arguments | What it does |
|---|---|---|
status | Recording state (recording, paused, stopping, stopped), record mode, and current problems. | |
control | action: start, pause, resume, stop | Control recording. Starting records the room’s microphone. |
set_mode | mode: always, meetings | Set the record mode. Places override it while you’re at one. |
Meetings and lines
| Tool | Arguments | What it does |
|---|---|---|
list_meetings | Past meetings, newest first: id, start, minutes, lines, first words. | |
delete_meeting | id | Delete a meeting with all its lines, fixes, tags and labels. |
list_lines | meeting, since, until, query, limit (all optional) | Lines in time order, as JSON. limit keeps the latest (default 200). |
get_line | id | One line. |
create_line | text, optional speaker (default “Note”), t (unix seconds, default now) | Add a note. A time inside a meeting puts it in that meeting. |
update_line | id, text and/or speaker | Fix text (teaches the transcriber) and/or set who said it (retrains voiceprints, a few seconds). Empty clears. |
delete_lines | ids | Delete lines with their fixes, tags and labels. All ids must exist. |
People
| Tool | Arguments | What it does |
|---|---|---|
list_people | Lines tagged per name, and the voiceprints with sample counts. Add a person by tagging a line with update_line. | |
rename_person | from, to | Move every line tagged from to to, then retrain. |
delete_person | name | Clear every tag with that name here and retrain. Ignored stops ignoring every ignored voice. |
Places and vocabulary
| Tool | Arguments | What it does |
|---|---|---|
list_places | Places and their settings. | |
set_place | label, action (record, meetings, off), optional lat, lon, radius (meters, default 150) | Create a place, or replace the one with that label. Without lat/lon, the menu bar sets it from the Mac’s location. |
delete_place | label | Delete a place. |
list_vocab | Vocabulary words. | |
add_vocab | words | Add words; existing ones are skipped. |
remove_vocab | words | Remove words. |
The menu bar app picks up changes to places, vocabulary and mode live. Argument details are in the schema each tool reports, generated from src/mcp.rs.
Files and data
Everything ozen keeps is in your ozen folder, ~/ozen, on your Mac. The only thing that leaves it is the voices
registry, which you host.
| Path | What it holds | Shared? |
|---|---|---|
transcript.txt | The readable transcript, one line per utterance | No |
lines.jsonl | Every line with its id, time, source, speaker guess, text and voiceprint | No |
tags.json | Speakers you set by tagging | No |
fixes.json | Text corrections you made | No |
learned.json | Hint words and automatic corrections learned from fixes | No |
fixes/ | Audio of fixed lines with the right text, for future fine-tuning | No |
ignore.json | Voiceprints of ignored voices | No |
vocab.txt | Terms the transcriber should spell right; edit freely | No (in the ozen repo) |
places.json | Your places, including where you live and work | No |
voices/ | The voices registry: names, voiceprints, accuracy history | Yes, with your other Macs through its git remote |
context/ | Folders written for AI agents | No |
chunks/ | Audio waiting to be transcribed; deleted once transcribed | No |
recent/ | The last 20 transcribed chunks, for eval.py. Set OZEN_KEEP_AUDIO=0 to keep none | No |
screen.png, screen-small.png | Last ozen look screenshot | No |
start.log | Logs from the recorder and transcriber | No |
Personal files (places.json, ignore.json, fixes*, context/) are gitignored in the ozen repo, so they
aren’t committed by accident.
Deleting data
- A meeting:
delete_meetingover MCP, which removes its lines, fixes and tags. - Specific lines:
delete_linesover MCP. - A person on this Mac:
delete_personover MCP. Remove them from the registry too, or other Macs keep them. - Everything: stop ozen and delete
~/ozen, and the registry repository if you no longer need it.
Tuning how fixes teach
Two settings decide how ozen learns from your fixes: how many learned words are hinted to the
transcriber, and how many times a correction must repeat before ozen applies it by itself. ozen eval scores
combinations of both so you can pick the best.
ozen eval --vocab 0,10,30 --repeat 0,1,2
It learns from half of a fixed set of spoken lines (as if you had fixed them) and scores each setting on the other half, best first:
- word error rate,
- English terms spelled right,
- word error rate on plain Hebrew,
- words invented on quiet noise.
By default the lines are synthetic (eval/cases.jsonl, spoken by macOS’s Hebrew voice). --real uses your own
fixes instead. --fresh ignores cached results.
Runs are deterministic and cached, so a rerun prints the same table and a sweep only transcribes new settings.
The synthetic set is small and has one voice, so treat small gaps as noise and confirm a winner with --real once
you have a few dozen fixes. To adopt a winner, set it as LEARN in
src/fixes.rs and rebuild.
Comparing models on your speech
From ~/ozen, uv run eval.py [N] transcribes your last N real chunks with the stock, Hebrew, and Hebrew plus
vocabulary setups side by side.
Troubleshooting
Start with the warnings at the top of the panel, or run ozen health: it prints one line per problem and nothing
when all is well. Logs are in ~/ozen/start.log.
“Recording blocked”
Ozen lacks Screen & System Audio Recording permission. Open System Settings › Privacy & Security › Screen & System Audio Recording, turn on Ozen, then press Start again. Do the same under Microphone if the room isn’t heard.
“Recording on hold: the screen is asleep or locked”
Screen capture, and with it call audio, pauses while the display sleeps or is locked. Recording resumes when you wake the screen.
“Microphone is silent”
The selected input device gives no sound. Pick another under System Settings › Sound › Input.
If ozen says it’s using one input because another is silent, it switched to a working microphone for you. Pick an input in System Settings to override.
“Transcriber stopped” or “Transcriber catching up”
- Stopped: the transcriber crashed. ozen restarts it automatically; queued audio is kept. If it keeps
happening, check
start.log. - Catching up: it’s more than about a minute behind. Normal on the first run while models download, or on a busy Mac. Lines will arrive late but none are lost.
No lines from the call
- Is the meeting app one ozen listens to (Zoom, Chrome, Teams, Slack, FaceTime, Discord)? Audio from other apps, like Safari, counts as computer audio and isn’t transcribed.
- In Meetings mode, ozen starts only once the meeting app opens the microphone. Press Start to record anyway.
Wrong speaker names
Tag a few lines per person and answer Review; each tag relabels the whole meeting. See Who is speaking.
A video or the computer’s voice shows up as a person
Click its speaker name and choose Ignore this voice. macOS Speak Selection in particular isn’t recognized as computer audio; see Limits.
Permissions reset after an update
Rebuild with cargo run --release -- app rather than copying the app by hand; the build signs it with the same
local certificate each time, which is what keeps permissions.
Limits
- Latency. Lines arrive about 15–30 seconds after speech. ozen transcribes in chunks, not word by word.
- Mixed language. English terms spoken inside Hebrew are the weakest spot. Add them to the vocabulary. Distant voices in the room are hard to hear.
- Echo. If you talk over the computer’s voice or a remote speaker and your voices sound alike, your line can be dropped as echo.
- Speak Selection. macOS Speak Selection isn’t captured as computer audio, so text it reads aloud is transcribed as a room speaker. Ignore this voice can drop it.
- Overlapping speech. Up to two voices at once are separated; a third merges into one of them. Overlaps in utterances under about 2 seconds aren’t detected, and utterances under 1 second take the previous speaker.
- Speaker matching is live, with no re-clustering after the meeting. Tagging a few lines fixes past and future labels.
- Platform. macOS 15+ on Apple Silicon only.
Privacy
ozen is built so meeting audio never leaves your Mac: no bot joins the call, and transcription and speaker recognition run locally. You’re still recording people, so a few things are on you.
Recording others
- ozen records and transcribes other people. Tell participants, and follow your local recording laws. Some places require everyone’s consent.
- In Always mode the mic records everything said near the Mac, not only meetings. Use Meetings mode or places to limit it.
Voiceprints are biometric data
- Keep the voices registry repository private.
- Enroll only people who agreed. Remove someone with
delete_person(see MCP server) and from the registry. - Ignored voices never go to the registry.
What leaves your Mac
| What | Where | When |
|---|---|---|
| Names and voiceprints | Your voices registry git remote | After tagging, and when pulling updates |
| Map tiles for the area shown | OpenStreetMap | Only while the Places map is open |
| Model downloads | Hugging Face and similar | First use of each model |
| Transcripts you hand to an agent | That agent’s AI provider | When you use Ask AI, Open, or MCP |
Your location, recordings, transcripts, fixes and screenshots otherwise stay on the Mac. Remember that the last row means cloud agents like Claude Code send what they read to their provider.
places.json holds where you live and work. Don’t copy it into shared folders.