Product & UX · 4 min read

CoCap

A native macOS app that turns large local video files into editable captions, without uploading anything.

CoCap at 45 percent, with a caption showing on the video and one caption line open for editing in the list on the right
Transcription is 45% done. The first minute of captions is already in the player and open for editing.

TL;DR

Problem
I needed realistic caption samples to test SubLime Captions. My own footage ran 20 to 40 GB per file, and in the browser a 30 GB upload took over an hour while long audio crashed the tab.
Solution
A native Mac app that transcribes on device, 30 seconds at a time, and puts each finished slice of captions into the player while the rest is still running.
Result
Released free and open source on GitHub. Memory stays around 150 MB whatever the file size, and nothing leaves the Mac.
View on GitHub ↗ (opens in a new tab)
Role
Solo product designer and builder: product strategy, interaction and SwiftUI implementation
Timeline
January to July 2026
Platform
macOS 14+ on Apple silicon
Tools
Swift, SwiftUI, Codex, Google AI Studio, Material Design 3
Status
Released on GitHub (v2.0.2)
  • ~150 MBof memory during transcription, for a 50 MB voice memo or a 40 GB camera master (Case 2)
  • 2.8 sto transcribe a 30-second window on an M3 Max, a real-time factor of 0.09 (Case 2)
  • 20–40 GBraw camera files and multi-track interviews that broke the browser version (Case 1)

CoCap started as a test utility for SubLime Captions and became a Mac app of its own. It transcribes local audio and video on device and gives you captions you can edit and export as SRT, VTT or TXT. I designed and built it alone, in SwiftUI with Codex. It runs on a 40 GB camera file with about 150 MB of memory, and the media never leaves the computer.

Context

While building SubLime Captions, a web tool that proofreads and translates subtitles, I kept needing large batches of subtitle samples with realistic speech-recognition errors. Converting files and collecting transcripts one by one was slow, so I decided to build a small tool that could make them on demand.

I made the first prototype in Google AI Studio. Giving the model speaker names and topic tags noticeably cut transcription errors. That was promising enough to move the project into Codex, add real file handling, and try it on longer media.

The first browser version, AI Media Captions, with a dashed drop zone for media and subtitle files
The first version ran in the browser. Short test clips worked fine.

Case 1 · I rebuilt CoCap as a native Mac app when real footage broke the browser

Question
How do I get files of 20 to 40 GB into a caption tool?
Options
Keep the web app and write chunked-upload scripts to keep the browser alive; or rebuild CoCap as a native macOS app that reads files from the local disk.
Trade-off
Chunked uploads would still send unreleased footage through a network and a cloud API, with the wait, the privacy question and the cost that come with it.
Decision
A native macOS app. When media stays on the local disk, the bandwidth and privacy problems go away.

Short clips worked in the browser. Then I tested with footage I had filmed before: raw 4K camera files and multi-track interviews that often weighed 20 to 40 GB. At that size the browser failed almost everywhere.

Uploading a 30 GB file on home broadband took over an hour, and a dropped connection meant starting again. When the browser tried to process long audio itself, the tab ran out of memory and crashed. Unreleased footage would have sat on a third-party cloud, and running commercial cloud APIs on hours of high-definition video got expensive fast.

CoCap now only goes online once, to download the speech model when you set it up. After that, transcription, editing and export all work offline, and media, captions and exports are never uploaded. It has no account system, analytics or telemetry.

Case 2 · Captions reach the player before the file is finished

Question
How does a laptop transcribe a multi-hour recording without running out of memory?
Options
Load the whole audio track and transcribe it in one pass; or stream the file in 30-second windows and show each result as soon as it is ready.
Trade-off
Loading everything at once makes memory grow with the file, and the editor sees nothing until the whole file is done.
Decision
A 30-second streaming pipeline. Memory stays flat, and finished captions go straight into the player.

Moving offline removed the network bottleneck and brought in a new limit: the computer’s memory. I benchmarked several speech models on Apple silicon, comparing transcription quality on Mandarin and English audio against memory use and launch time. I chose an on-device model that runs on the Mac’s GPU through Apple’s Metal Performance Shaders. On an M3 Max it transcribes a 30-second window in 2.8 seconds, an average real-time factor of 0.09.

Thirty seconds at a time

CoCap reads the video with Apple’s own AVFoundation decoders and holds only the current 30-second slice of audio plus the next one in memory. Memory stays around 150 MB whether the input is a 50 MB voice memo or a 40 GB camera master.

A flow from a local media file to a 30-second audio buffer to the local MPS engine, which feeds both the video player preview and the subtitle editor
Each 30-second window goes to the on-device engine, and its captions go to the player and the editor at the same time.

Streaming also means editors don’t have to wait. Most transcription software makes you wait for the whole file before it shows any text. CoCap puts each finished slice into the video player right away, so an editor can scrub back and check the first scene while the rest of the file is still transcribing.

Cuts where people breathe

People rarely speak in neat sentences, and a hard cut every 30 seconds would split words in half. CoCap uses voice activity detection to find pauses and overlaps neighbouring chunks slightly, so syllables aren’t clipped and caption breaks land where the speaker takes a breath.

Pause keeps what’s done

Stopping a long job shouldn’t throw away finished work. Pause shows “Pausing transcription…” while the engine finishes its current slice. After that, every completed caption stays editable and can be exported while the rest waits.

Case 3 · Three punctuation choices replace the model settings

Question
What should a video editor be able to adjust?
Options
Expose the speech model's settings, such as temperature and beam width; or offer only the choice editors make anyway, how punctuation looks on screen.
Trade-off
Editors rarely know what the model settings do, and the result they care about is the text on screen.
Decision
Three plain-language punctuation options. The model's mechanics stay hidden.

Keep All keeps full punctuation, for narrative scripts. Remove Trailing drops the full stop or comma at the end of each caption, following broadcast subtitle conventions. Remove All gives unpunctuated text for fast social clips. A one-line description under the menu says what the current choice does.

The Caption punctuation menu open below the video, with Keep All, Remove Trailing (selected) and Remove All
The only transcription setting in the workbench. Each option names the result on screen.

A Mac app built on Material Design 3

Writing interface code with an AI assistant tends to drift: padding and colours change from one view to the next. I gave Codex Google’s Material Design 3 as its guide. MD3 has explicit rules for type scale, surface elevation and tonal colour roles, and those rules kept the generated SwiftUI consistent from one revision to the next.

I adapted the MD3 surfaces to feel at home on macOS. CoCap uses a unified title bar, the system file pickers and standard shortcuts, such as Space to play. The layout is one split view: the video and its details on the left, the caption document on the right.

First-time setup happens in that same window. If the speech model is missing at launch, the media panel shows a download card with a progress bar, and every file is checked against its checksum before it is installed. When the download finishes, the drop zone takes the card’s place, so there is no separate setup wizard.

The empty workbench: a dashed drop zone on the left listing supported formats, and an empty caption panel on the right
Ready to work: the drop zone takes MP4, MOV, MKV, WebM, TS, MP3 and WAV.

Settings that are not part of the editing work live in a standard macOS settings window: appearance, interface language and the state of the local model.

The CoCap Settings window with appearance, interface language, the local transcription model and a Model ready status
The settings window follows macOS conventions.

How I built it

I didn’t draw mockups in Figma for CoCap. Working alone, static screens felt like extra overhead when I could change the SwiftUI code directly with Codex and try the result on a real file.

I wrote the code with Codex. My part was deciding what to build and checking the result: choosing native over web, running the model benchmarks and picking the model, setting the 30-second memory limit, and deciding what editors see and what stays hidden. MD3 tokens gave Codex concrete limits for type, spacing and colour, so I could ask for fast changes without the interface coming apart.

Outcome

CoCap is released free on GitHub under the MIT licence, currently at version 2.0.2, for macOS 14 or later on Apple silicon. The first model download is about 5.6 GB. After that, the app works fully offline. The build is ad-hoc signed and not yet notarized by Apple, so macOS asks the user to approve it on first launch.

There is no usage data yet. CoCap has no analytics by design, so feedback will have to come from GitHub issues and from editors I can talk to directly.

Reflection

Pairing an AI coding assistant with an explicit design system worked much better for me than starting from visual mockups. The system gave the AI limits to work inside, and the running app gave me something real to judge.

The project also showed me the value of sensible defaults. Editors needed control over how captions look on screen, so I put the model’s settings out of sight and gave them three punctuation options.

What’s next

I want to ship a Developer ID signed and notarized build, so installing CoCap doesn’t need a trip to System Settings. Then I’d like to test it with video editors working on their own footage.