Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

ozen

ozen (אוזן, “ear”) is a local meeting copilot for macOS. While you’re in a Zoom, Meet or Teams call, ozen listens, transcribes the call and the room in Hebrew and English, knows who is speaking, and lets an AI assistant (Claude Code or Hermes) read the meeting and see your screen so it can answer questions as the meeting happens.

Everything runs on your Mac. No bot joins the meeting and no audio leaves the machine.

What you get

  • A live transcript of the call and of people in the room, one line per utterance, with a speaker name on every line. Lines appear about 15–30 seconds after they’re spoken.
  • Speaker names that stick. ozen recognizes voices it has heard before. Tell it once who said a line and it relabels the rest of the meeting, and remembers that person next time.
  • A transcript that learns. Fix a misheard word and ozen learns the word; fix the same mistake twice and it corrects it by itself from then on.
  • Echo and noise filtered out. Remote voices leaking from your speakers into the mic, the computer reading text aloud, and a video playing nearby don’t pollute the transcript.
  • Overlapping speech separated. When two people talk at once, each gets their own line.
  • Recording on your terms. Record always, only during meetings, or by place (auto on at the office, off at home).
  • AI on tap. Hand one or more past meetings, or the one happening right now, to an AI agent with one click. Any MCP-capable agent can also read and edit ozen’s data directly.

A transcript line looks like this:

[15:06:52] Ofek Gabay (room): אימבדין שלי, אני לא מוצא את זה
[15:09:41] S2 (call): Let's move to the launch timeline.

(call) lines came from the meeting app, (room) lines from your microphone. S2 is a voice ozen hasn’t been told a name for yet.

Where to start

  1. Install ozen and grant its permissions.
  2. Record your first meeting.
  3. Learn the menu bar panel, where you’ll spend most of your time.

Before recording other people, read Privacy.

Install

Requirements

  • A Mac with Apple Silicon running macOS 15 or later
  • Xcode command line tools (xcode-select --install), for swiftc
  • Rust (cargo)
  • uv (Python runner)
  • ffmpeg (brew install ffmpeg)
  • Read access to the ozen repository and to your voices registry

Build and install

git clone https://github.com/ozenhq/ozen ~/ozen
git clone https://github.com/tupe12334/voices-embedding-registry ~/ozen/voices
cd ~/ozen && cargo run --release -- app

The last command builds the ozen command-line tool (~/ozen/target/release/ozen) and installs Ozen.app into ~/Applications.

Open Ozen from Spotlight, Launchpad or Finder like any app. It lives in the menu bar as an ear icon and has no Dock icon.

Tip: put the CLI on your PATH so you can type ozen instead of ~/ozen/target/release/ozen: add alias ozen=~/ozen/target/release/ozen to ~/.zshrc. The rest of these docs write ozen.

Permissions

The first time you press Start, macOS asks Ozen for two permissions. Grant both in System Settings › Privacy & Security, then press Start again:

PermissionWhy
Screen & System Audio RecordingTo capture audio from meeting apps and other apps (ScreenCaptureKit). ozen does not record video.
MicrophoneTo hear people in the room, including you.
Location Services (optional)Asked only when you first locate a place. Without it, places never match and recording follows Always / Meetings.

The build signs the app with a local self-signed certificate (kept in ~/Library/Keychains/ozen-signing.keychain-db), so permissions survive rebuilds.

First-run downloads

The speech models download the first time they’re needed: Whisper (Hebrew and English) and the voice-recognition model on the first transcription, and the ~640 MB speech separator in the background after the first start. Until the separator is ready, two people talking at once stay merged in one line. While models download the panel may say the transcriber is catching up; that’s expected.

Updating

cd ~/ozen && git pull && cargo run --release -- app

Rebuilding never stops a recording in progress.

Another Mac

Run the same three install commands on the other Mac. Voices you tag on either Mac reach the other within a few minutes; see Using several Macs.

Your first meeting

  1. Start recording. Click the ear icon in the menu bar and press Start. The icon turns into a red filled ear while ozen records.
  2. Join your call in Zoom, Meet (in Chrome), Teams, Slack, FaceTime or Discord. Nothing to configure: ozen hears meeting apps directly, not through your speakers.
  3. Watch the transcript. Open the panel from the ear icon. New lines arrive every few seconds, about 15–30 seconds behind speech. Hebrew lines read right to left.
  4. Name the speakers. Voices ozen doesn’t know show as S1, S2, … Click a speaker name and pick or type who it is. ozen relabels that voice everywhere in the meeting and recognizes it next time. See Who is speaking.
  5. Fix a word if one was misheard: click the line’s text and type the correction. See Fixing the transcript.
  6. Ask AI. Press Ask AI to open Claude Code or Hermes on the meeting happening now. Ask “what did we decide about the launch date?” or “what’s on my screen?”. See Working with AI agents.
  7. Stop. Press Stop. ozen finishes transcribing what it has already heard (the icon shows an hourglass), then stops.

Don’t want to press Start every time?

Switch the panel to Meetings mode and ozen starts by itself whenever a meeting app uses the microphone, and stops shortly after the call ends. See When ozen records.

The menu bar panel

Ozen lives in the menu bar as an ear icon.

  • Click the icon for the panel.
  • Right-click for a menu with the same controls (Start / Pause / Resume / Stop, Always / Meetings, Places…).

The icon

IconMeaning
Red filled earRecording
Monochrome earStopped
Pause symbolPaused
HourglassStopped, finishing transcription of audio already heard

When not recording, the panel says Not recording and why (paused, finishing transcription, or a problem).

Controls

ButtonWhat it does
StartStart recording the call, the room microphone and other apps’ audio.
Pause / ResumeStop and restart listening. The transcriber stays loaded, so resuming is instant.
StopStop listening, finish transcribing queued audio (up to 2 minutes), then unload everything.
Always / MeetingsRecord mode; see When ozen records.
Places…Place-based rules; see Places.
Review NJump to the line ozen is least sure who said; see Review.
Ask AIOpen an AI agent on the meeting happening now; see Working with AI agents.

Problems show as warnings at the top of the panel: recording blocked by a missing permission, a silent microphone, a screen that’s asleep, or a transcriber that’s behind. See Troubleshooting.

Views

Transcript

The live transcript, refreshed every 2 seconds. Each line shows its time, speaker and text.

  • Click a speaker name to say who really said it, or to ignore that voice.
  • Click the text to correct it.
  • An orange ? marks a line ozen is unsure about.
  • Grey lines are ignored voices.

Timeline

One lane per speaker, with their total talk time, and a bar for every line they spoke along a scrollable time axis. Overlapping speech overlaps here too.

  • − / + zoom.
  • Silences over 2 minutes shrink to a short break marker.
  • Hover a bar to read the line; click it to jump to that line in the transcript.

Meetings

Every past meeting, newest first. A silence of 10 minutes or more starts a new meeting.

Select one or more (⌘-click or ⇧-click) and press Open to hand them to an AI agent; see Working with AI agents.

The footer shows speaker-recognition accuracy, now and when you started.

When ozen records

What ozen hears

While recording, ozen captures three audio streams:

StreamSourceTranscribed?
callMeeting apps only: Zoom, Chrome, Teams, Slack, FaceTime, DiscordYes, as (call) lines
micYour microphone, meaning everyone in the roomYes, as (room) lines
localEvery other app, like a video or the say commandNo. Used only to recognize when the computer itself is talking, so it isn’t mistaken for a person in the room

Record modes

Pick a mode with the Always / Meetings toggle in the panel or the right-click menu.

Always. Records from Start until you pause or stop, including everything the room mic hears.

Meetings. Starts by itself when a meeting app (Zoom, Chrome/Meet, Teams, Slack, FaceTime, Discord) starts using the microphone, and stops 20 seconds after it releases it. This works for any call in those apps, with no calendar or plugin needed.

In either mode, pausing or starting by hand wins until the next meeting starts or ends. Relaunching Ozen never stops a recording in progress.

Places

Places switch recording by where your Mac is. Open Places… from the panel or the right-click menu.

Each place has a label, a location, a radius, and one of three settings:

SettingWhile you’re there
Auto recordRecord (like Always)
Record meetings onlyRecord during calls (like Meetings)
Auto offDon’t record

While you’re inside a place’s radius, its setting replaces Always / Meetings. Arriving or leaving applies right away; a manual pause or start holds until then. Outside every place, your Always / Meetings choice applies.

Home and Work are there from the start with no location, so they do nothing until you set them. To set a place’s location:

  • type its latitude and longitude,
  • press Use current location,
  • press Pick on map and click, or
  • drag its pin on the map.

The first time, macOS asks for Location Services. Right after Ozen launches it waits for your location before any place decides anything.

The map uses OpenStreetMap; viewing it downloads map tiles for that area. Your location itself is never sent anywhere. Places are saved in ~/ozen/places.json; see Files and data.

Who is speaking

ozen gives every line a speaker by its voice. Each utterance gets a voiceprint, which is compared with voices already heard, so one person keeps one label for the whole meeting.

  • Known people show by name. They come from your voices registry, which grows as you tag.
  • Unknown voices show as S1, S2, … until you name them.

Naming a speaker (tagging)

Click any speaker name in the transcript and pick or type who really said that line. Then ozen:

  1. rebuilds that person’s voiceprint from every line tagged as them, in this meeting and earlier ones,
  2. relabels every untagged line with the new voiceprints, past and future,
  3. saves the voice to the registry, so your other Macs learn it too.

This takes a few seconds. A few tags per person early in a meeting fix most labels. To undo a tag, click the speaker name and choose Clear tag.

From the command line: ozen tag <line-id> "Dana Levi" (an empty name clears the tag).

Review

ozen scores every untagged line and queues the ones it’s least sure about: near the same-voice cutoff, or nearly tied between two people. They show an orange ? in the transcript.

Press Review N to jump to the most uncertain line and answer who said it. Review asks only about the last 10 minutes, because after that nobody remembers who said what.

Answering Review is the fastest way to improve recognition. In a test with four similar voices, 12 tags chosen by Review reached 100% accuracy on untagged lines, against 98% for 12 random tags.

Every tag also recalibrates how similar two clips must be to count as the same voice, and logs accuracy; the panel footer shows it next to the starting value.

Ignoring a voice

A video playing next to your Mac, or a radio in the room, isn’t part of the meeting. Click its speaker name and choose:

  • Ignore this voice to ignore that line’s voice, or
  • Ignore all N lines by S3 to mark every nearby line of that speaker at once.

Those lines, and earlier ones that sound like them, turn grey and leave the timeline, and ozen stops writing that voice from then on. Ignored voices stay on this Mac; they’re never shared to the registry.

To undo, tag the line as a person, or choose Clear tag.

From the command line: ozen ignore <line-id>....

People talking at once

When two people talk at once, in the room, on the call, or one on each, ozen separates the audio into one track per voice and writes each as its own line with its own speaker and time. In the timeline, those lines overlap.

Only utterances where voices disagree are separated, so a single speaker costs nothing extra. See Limits for what isn’t separated.

Echo

A mic utterance that mostly overlaps call or computer audio, in the same voice, is your speakers bleeding into the mic, not a person in the room, so ozen drops it. That covers the computer reading text aloud and remote voices leaking into the mic.

Someone in the room talking over the call keeps their line, because their voice doesn’t match what was playing.

Fixing the transcript

Correcting a line

Click a line’s text in the panel and type what was really said. Leave it empty to restore what ozen heard.

From the command line: ozen fix <line-id> "right text" (empty clears).

What ozen learns from fixes

  • New words. Words your fixes add, like names, product terms and jargon, are hinted to the transcriber on every new line, most used first.
  • Repeated corrections. Correct the same phrase the same way twice (say “פי אר” → “PR”) and ozen applies that correction to new lines automatically.
  • A corrected line keeps what the transcriber originally heard, so if an automatic correction is wrong, fixing it back cancels it. A correction is also skipped whenever one of your fixes kept that phrase as right.

Each fix also keeps its audio with the right text, as a dataset for fine-tuning a model later. It stays on your Mac.

Vocabulary

~/ozen/vocab.txt lists names and terms the transcriber should spell right, like Kev, PR, code review. Edit it freely, one term per line or comma-separated; it’s read on every line. Known people’s names are hinted automatically.

Agents can manage it too, with the list_vocab, add_vocab and remove_vocab MCP tools.

Languages

Each utterance is transcribed as either Hebrew or English, picked per utterance. Hebrew uses a Hebrew-trained Whisper model (ivrit.ai), English uses stock Whisper large-v3-turbo. English terms spoken inside Hebrew are the weakest spot; add them to the vocabulary.

Filler that Whisper tends to invent on silence (“Thank you.”, “תודה רבה”) is dropped.

Using several Macs

Voices are shared across your Macs through the voices registry, a private git repository that every Mac clones into ~/ozen/voices.

  • Start pulls the registry.
  • While running, ozen pulls again every few minutes.
  • Every tag rebuilds voiceprints from the newest registry and pushes when done.

So a person you tag on one Mac is recognized on the other within a few minutes.

Tagging on two Macs at once is safe. A push that loses the race rebuilds on top of the other Mac’s and pushes again, so both sets of tags survive.

Offline tags are committed locally and pushed with the next tag once you’re back online.

What stays on each Mac

Each Mac’s own tags and transcript never leave it; the registry holds only names and voiceprints. Ignored voices, places, fixes and recordings are local too. See Files and data.

Don’t edit ~/ozen/voices by hand. Every retrain resets it to the remote before rebuilding, so hand edits are lost. Change voices by tagging.

Setting up another Mac

Follow Install on the new Mac, with the same registry URL.

Working with AI agents

ozen hands meetings to an AI coding agent, Claude Code or Hermes, in a folder holding the transcripts plus an AGENTS.md (and CLAUDE.md) that tells the agent what’s there. The agent must be installed and on your PATH.

The meeting happening now

Press Ask AI in the panel. ozen writes the current meeting (one whose last line is under 10 minutes old) to ~/ozen/context/live/ and starts the agent there in a new Terminal window.

While the meeting goes on, a background process rewrites that folder with the latest lines every 15 seconds, and exits when the meeting ends. The agent is told to reread it, and that ozen look shows your screen, so you can ask things like:

  • “Summarize the last five minutes.”
  • “What did Dana ask me to do?”
  • “What’s on my screen right now, and does it match what they’re describing?”

From the command line: ozen live --open claude (or hermes); without --open it only writes the folder and prints its path.

Past meetings

Switch the panel to Meetings, select one or more (⌘-click or ⇧-click), and press Open. ozen writes their transcripts into a fresh folder under ~/ozen/context/ and starts the agent you choose there. The folder’s claude.command or hermes.command reopens it later with a double-click.

Auto add with Kev also adds every other meeting that a local Kev server (localhost:8009) judges part of the same project or topic. Its scores show before you pick the agent. Without Kev running, this option does nothing.

From the command line:

ozen meetings                # list meetings with their ids
ozen gather [--kev] ID...    # write them to context/<now>/ and print the folder
ozen open DIR claude         # or hermes, or finder

Screen and recent lines

ozen look [N] takes a screenshot of the main display (saved as ~/ozen/screen-small.png) and prints the last N transcript lines (default 40), with tagged speakers and fixed text. Agents use it to answer “what’s on screen” or “what was just said”. The app running it (like your terminal) needs Screen Recording permission.

Any MCP agent

ozen mcp gives any MCP client full access to meetings, transcript lines, people, places, vocabulary and recording control. See MCP server.

Command line

The ozen binary is built to ~/ozen/target/release/ozen. It works from any directory. Most commands do what a panel button does, so you can script ozen or drive it from an agent.

Recording

CommandWhat it does
ozen startStart recording (same as Start).
ozen pausePause; the transcriber stays loaded.
ozen resumeResume after a pause.
ozen stopStop, finish transcribing queued audio (up to 2 minutes), then unload.
ozen statusPrint recording, paused, stopping or stopped.
ozen healthPrint one line per problem the panel warns about; nothing when all is well. See Troubleshooting.

Transcript

CommandWhat it does
ozen show [N]Last N lines, with tagged speakers and fixed text.
ozen look [N]Screenshot to ~/ozen/screen-small.png, then the last N lines (default 40).
ozen fix <line-id> "right text"Correct a line’s text; empty restores it. Relearns hint words and corrections.
ozen tag <line-id> "Name"Say who said a line; empty clears. Retrains and pushes the registry.
ozen ignore <line-id>...Mark lines as a voice to ignore. Retrains.
ozen retrainRebuild voiceprints, labels, ignored voices and accuracy from all tags. Tagging does this for you.

Meetings and agents

CommandWhat it does
ozen meetingsPast meetings, newest first: id, start, minutes, lines, first words.
ozen gather [--kev] ID...Write those meetings to ~/ozen/context/<now>/ and print the folder. --kev adds meetings Kev judges related.
ozen live [--open claude|hermes]Write the current meeting to ~/ozen/context/live/, keep it updated every 15 s until the meeting ends, and print the folder. --open starts that agent there.
ozen open DIR claude|hermes|finderOpen a folder written by gather or live in that agent, or in Finder.
ozen mcpRun the MCP server on stdio.

App

CommandWhat it does
ozen appBuild and install ~/Applications/Ozen.app.
ozen barBuild if needed, then open Ozen.app.

Evaluation

CommandWhat it does
ozen eval [--vocab 0,10,30] [--repeat 0,1,2] [--real] [--fresh]Score learning settings; see Tuning how fixes teach.
uv run eval.py [N]Transcribe your last N real chunks with stock, Hebrew and Hebrew+vocabulary models, to compare on your own speech. Run from ~/ozen.

MCP server

ozen mcp is a Model Context Protocol server on stdio. It lets any MCP-capable agent read and edit what ozen keeps: meetings, transcript lines, people, places, vocabulary, and recording.

Connect it

Claude Code:

claude mcp add -s user ozen -- ~/ozen/target/release/ozen mcp

Any other MCP client: command ~/ozen/target/release/ozen, argument mcp.

Concepts

  • Lines are transcript lines: id, time, speaker, text. A line’s speaker is either a tag (set by you or an agent), which wins, or the voiceprint’s guess. text is the fixed text if the line was fixed; heard is what the transcriber originally wrote.
  • Meetings are runs of lines with under 10 minutes of silence between them. A meeting’s id is its start time in unix seconds.
  • Notes are lines an agent creates, like a summary or an action item. They have no voice, so editing one just rewrites it.
  • The speaker name Ignored marks a voice to drop.

Tools

Recording

ToolArgumentsWhat it does
statusRecording state (recording, paused, stopping, stopped), record mode, and current problems.
controlaction: start, pause, resume, stopControl recording. Starting records the room’s microphone.
set_modemode: always, meetingsSet the record mode. Places override it while you’re at one.

Meetings and lines

ToolArgumentsWhat it does
list_meetingsPast meetings, newest first: id, start, minutes, lines, first words.
delete_meetingidDelete a meeting with all its lines, fixes, tags and labels.
list_linesmeeting, since, until, query, limit (all optional)Lines in time order, as JSON. limit keeps the latest (default 200).
get_lineidOne line.
create_linetext, optional speaker (default “Note”), t (unix seconds, default now)Add a note. A time inside a meeting puts it in that meeting.
update_lineid, text and/or speakerFix text (teaches the transcriber) and/or set who said it (retrains voiceprints, a few seconds). Empty clears.
delete_linesidsDelete lines with their fixes, tags and labels. All ids must exist.

People

ToolArgumentsWhat it does
list_peopleLines tagged per name, and the voiceprints with sample counts. Add a person by tagging a line with update_line.
rename_personfrom, toMove every line tagged from to to, then retrain.
delete_personnameClear every tag with that name here and retrain. Ignored stops ignoring every ignored voice.

Places and vocabulary

ToolArgumentsWhat it does
list_placesPlaces and their settings.
set_placelabel, action (record, meetings, off), optional lat, lon, radius (meters, default 150)Create a place, or replace the one with that label. Without lat/lon, the menu bar sets it from the Mac’s location.
delete_placelabelDelete a place.
list_vocabVocabulary words.
add_vocabwordsAdd words; existing ones are skipped.
remove_vocabwordsRemove words.

The menu bar app picks up changes to places, vocabulary and mode live. Argument details are in the schema each tool reports, generated from src/mcp.rs.

Files and data

Everything ozen keeps is in your ozen folder, ~/ozen, on your Mac. The only thing that leaves it is the voices registry, which you host.

PathWhat it holdsShared?
transcript.txtThe readable transcript, one line per utteranceNo
lines.jsonlEvery line with its id, time, source, speaker guess, text and voiceprintNo
tags.jsonSpeakers you set by taggingNo
fixes.jsonText corrections you madeNo
learned.jsonHint words and automatic corrections learned from fixesNo
fixes/Audio of fixed lines with the right text, for future fine-tuningNo
ignore.jsonVoiceprints of ignored voicesNo
vocab.txtTerms the transcriber should spell right; edit freelyNo (in the ozen repo)
places.jsonYour places, including where you live and workNo
voices/The voices registry: names, voiceprints, accuracy historyYes, with your other Macs through its git remote
context/Folders written for AI agentsNo
chunks/Audio waiting to be transcribed; deleted once transcribedNo
recent/The last 20 transcribed chunks, for eval.py. Set OZEN_KEEP_AUDIO=0 to keep noneNo
screen.png, screen-small.pngLast ozen look screenshotNo
start.logLogs from the recorder and transcriberNo

Personal files (places.json, ignore.json, fixes*, context/) are gitignored in the ozen repo, so they aren’t committed by accident.

Deleting data

  • A meeting: delete_meeting over MCP, which removes its lines, fixes and tags.
  • Specific lines: delete_lines over MCP.
  • A person on this Mac: delete_person over MCP. Remove them from the registry too, or other Macs keep them.
  • Everything: stop ozen and delete ~/ozen, and the registry repository if you no longer need it.

Tuning how fixes teach

Two settings decide how ozen learns from your fixes: how many learned words are hinted to the transcriber, and how many times a correction must repeat before ozen applies it by itself. ozen eval scores combinations of both so you can pick the best.

ozen eval --vocab 0,10,30 --repeat 0,1,2

It learns from half of a fixed set of spoken lines (as if you had fixed them) and scores each setting on the other half, best first:

  • word error rate,
  • English terms spelled right,
  • word error rate on plain Hebrew,
  • words invented on quiet noise.

By default the lines are synthetic (eval/cases.jsonl, spoken by macOS’s Hebrew voice). --real uses your own fixes instead. --fresh ignores cached results.

Runs are deterministic and cached, so a rerun prints the same table and a sweep only transcribes new settings.

The synthetic set is small and has one voice, so treat small gaps as noise and confirm a winner with --real once you have a few dozen fixes. To adopt a winner, set it as LEARN in src/fixes.rs and rebuild.

Comparing models on your speech

From ~/ozen, uv run eval.py [N] transcribes your last N real chunks with the stock, Hebrew, and Hebrew plus vocabulary setups side by side.

Troubleshooting

Start with the warnings at the top of the panel, or run ozen health: it prints one line per problem and nothing when all is well. Logs are in ~/ozen/start.log.

“Recording blocked”

Ozen lacks Screen & System Audio Recording permission. Open System Settings › Privacy & Security › Screen & System Audio Recording, turn on Ozen, then press Start again. Do the same under Microphone if the room isn’t heard.

“Recording on hold: the screen is asleep or locked”

Screen capture, and with it call audio, pauses while the display sleeps or is locked. Recording resumes when you wake the screen.

“Microphone is silent”

The selected input device gives no sound. Pick another under System Settings › Sound › Input.

If ozen says it’s using one input because another is silent, it switched to a working microphone for you. Pick an input in System Settings to override.

“Transcriber stopped” or “Transcriber catching up”

  • Stopped: the transcriber crashed. ozen restarts it automatically; queued audio is kept. If it keeps happening, check start.log.
  • Catching up: it’s more than about a minute behind. Normal on the first run while models download, or on a busy Mac. Lines will arrive late but none are lost.

No lines from the call

  • Is the meeting app one ozen listens to (Zoom, Chrome, Teams, Slack, FaceTime, Discord)? Audio from other apps, like Safari, counts as computer audio and isn’t transcribed.
  • In Meetings mode, ozen starts only once the meeting app opens the microphone. Press Start to record anyway.

Wrong speaker names

Tag a few lines per person and answer Review; each tag relabels the whole meeting. See Who is speaking.

A video or the computer’s voice shows up as a person

Click its speaker name and choose Ignore this voice. macOS Speak Selection in particular isn’t recognized as computer audio; see Limits.

Permissions reset after an update

Rebuild with cargo run --release -- app rather than copying the app by hand; the build signs it with the same local certificate each time, which is what keeps permissions.

Limits

  • Latency. Lines arrive about 15–30 seconds after speech. ozen transcribes in chunks, not word by word.
  • Mixed language. English terms spoken inside Hebrew are the weakest spot. Add them to the vocabulary. Distant voices in the room are hard to hear.
  • Echo. If you talk over the computer’s voice or a remote speaker and your voices sound alike, your line can be dropped as echo.
  • Speak Selection. macOS Speak Selection isn’t captured as computer audio, so text it reads aloud is transcribed as a room speaker. Ignore this voice can drop it.
  • Overlapping speech. Up to two voices at once are separated; a third merges into one of them. Overlaps in utterances under about 2 seconds aren’t detected, and utterances under 1 second take the previous speaker.
  • Speaker matching is live, with no re-clustering after the meeting. Tagging a few lines fixes past and future labels.
  • Platform. macOS 15+ on Apple Silicon only.

Privacy

ozen is built so meeting audio never leaves your Mac: no bot joins the call, and transcription and speaker recognition run locally. You’re still recording people, so a few things are on you.

Recording others

  • ozen records and transcribes other people. Tell participants, and follow your local recording laws. Some places require everyone’s consent.
  • In Always mode the mic records everything said near the Mac, not only meetings. Use Meetings mode or places to limit it.

Voiceprints are biometric data

  • Keep the voices registry repository private.
  • Enroll only people who agreed. Remove someone with delete_person (see MCP server) and from the registry.
  • Ignored voices never go to the registry.

What leaves your Mac

WhatWhereWhen
Names and voiceprintsYour voices registry git remoteAfter tagging, and when pulling updates
Map tiles for the area shownOpenStreetMapOnly while the Places map is open
Model downloadsHugging Face and similarFirst use of each model
Transcripts you hand to an agentThat agent’s AI providerWhen you use Ask AI, Open, or MCP

Your location, recordings, transcripts, fixes and screenshots otherwise stay on the Mac. Remember that the last row means cloud agents like Claude Code send what they read to their provider.

places.json holds where you live and work. Don’t copy it into shared folders.