Transcribe audio on iOS & macOS: WhisperKit
WhisperKit runs Whisper speech recognition models on Apple devices. This guide pins the Swift
package to 0.9.0 and provides a complete macOS command-line example for transcribing audio files.
The same library supports iOS apps; their build and permission requirements are described separately.
Introduction to WhisperKit and its capabilities
WhisperKit uses Core ML for local inference. Its initial setup can download models and tokenizer files from Hugging Face. Local processing does not mean that the first run is disconnected: provision and validate those assets before using the app without network access.
Setting up WhisperKit on iOS and macOS
To integrate WhisperKit into your apps, you'll need to set up the proper environment:
Prerequisites
- An Apple Silicon Mac for the native example in this guide.
- Swift
5.9or later. Apple Command Line Tools can build this macOS Swift package; full Xcode is required to build and run an iOS app. - Deployment targets of iOS
16or macOS13and later, as declared by the pinned package manifest. - A local audio file and enough storage for the selected model and tokenizer assets.
The example is verified on macOS. That verification does not establish an iOS simulator or device build, support for every deployment target, or transcription speed on other hardware.
Installation
Create an empty directory containing this Package.swift. The package is pinned rather than using
an open-ended minimum version, so the API contract matches the examples below.
// swift-tools-version: 5.9
import PackageDescription
let package = Package(
name: "AudioTranscription",
platforms: [.iOS(.v16), .macOS(.v13)],
dependencies: [
.package(url: "https://github.com/argmaxinc/WhisperKit.git", exact: "0.9.0")
],
targets: [
.executableTarget(name: "TranscribeAudio", dependencies: [
.product(name: "WhisperKit", package: "WhisperKit")
])
]
)
For an iOS app, add that repository and exact version through Xcode’s package dependency interface,
select the WhisperKit library product, and set the app’s deployment target to at least iOS 16.
The command-line entry point below is for macOS; invoke the library from your app’s own task or
view model on iOS.
Step-by-step guide to transcribing audio files
Save this complete program as Sources/TranscribeAudio/TranscribeAudio.swift. Pass an audio path
and a writable model cache directory. An optional third argument supplies an already downloaded
model directory. This file-based example does not request microphone permission.
import Foundation
import WhisperKit
@main
struct TranscribeAudio {
static func main() async {
let arguments = CommandLine.arguments
guard arguments.count == 3 || arguments.count == 4,
FileManager.default.isReadableFile(atPath: arguments[1]) else {
FileHandle.standardError.write(Data(
"Usage: TranscribeAudio readable-audio-file cache-directory [model-directory]\n".utf8
))
exit(1)
}
do {
let cache = URL(fileURLWithPath: arguments[2], isDirectory: true)
try FileManager.default.createDirectory(at: cache, withIntermediateDirectories: true)
let localModel = arguments.count == 4 ? arguments[3] : nil
let config = WhisperKitConfig(
model: "tiny.en",
downloadBase: cache,
modelRepo: "argmaxinc/whisperkit-coreml",
modelFolder: localModel,
tokenizerFolder: cache,
verbose: false,
load: true,
download: localModel == nil
)
let pipe = try await WhisperKit(config)
let options = DecodingOptions(
language: "en",
skipSpecialTokens: true,
concurrentWorkerCount: 1,
chunkingStrategy: .vad
)
let results: [TranscriptionResult] = try await pipe.transcribe(
audioPath: arguments[1], decodeOptions: options
)
print(results.map(\.text).joined(separator: "\n"))
} catch {
FileHandle.standardError.write(Data("Audio transcription failed.\n".utf8))
exit(1)
}
}
}
swift build --product TranscribeAudio
swift run --skip-build TranscribeAudio "recording.wav" Models
The explicit [TranscriptionResult] type selects the array-returning API. Join every result in
order: the deprecated optional-result overload returns only the first result and can discard later
chunks. See the pinned transcription implementation.
The audio loader uses AVFoundation. Start with a supported local audio file such as PCM WAV, MP3,
or AAC in M4A. For video, extract the soundtrack first.
In an app, retain and reuse a loaded WhisperKit instance and serialize access to it instead of
loading the models for every recording.
Using a specific model
The program explicitly chooses tiny.en, an English-only model. Change model to base.en for
another English-only model, or use a multilingual variant such as base. For a different spoken
language, set the corresponding DecodingOptions.language value too. Check the actual Core ML
folder names in the model repository
before changing the configuration; the English suffix uses a period, as in tiny.en.
Optimizing transcription accuracy and performance
To achieve optimal results with WhisperKit:
Audio quality considerations
- Use clear, high-quality audio recordings when possible.
- Minimize background noise in recording environments.
- For voice recordings, position microphones closer to speakers.
Model selection
The model families include tiny.en, base.en, small.en, and medium.en for English, with
multilingual variants such as tiny, base, and large-v3. Their Core ML exports and compression
variants have different storage and runtime costs. Measure accuracy and memory on representative
recordings and target devices before selecting a larger model.
Performance considerations
- Download size does not equal peak inference memory; account for the model, decoder state, and audio.
- Processing time depends on the model, audio length, and target hardware.
- Start with a smaller model and reuse the loaded pipeline.
Practical use cases and examples
WhisperKit can be effectively used in:
- Voice note applications with automatic transcription.
- Accessibility features for hearing-impaired users.
- Meeting and interview transcription tools.
- Podcast and video content transcription.
- Language learning applications.
Troubleshooting common issues
Model download failures
Issue: Models fail to download or initialize.
Solution: Check network access, available storage, and the selected model name. After a successful first run, reuse the downloaded model with the program’s third argument:
swift run --skip-build TranscribeAudio "recording.wav" Models \
Models/models/argmaxinc/whisperkit-coreml/openai_whisper-tiny.en
modelFolder must point to the directory containing AudioEncoder.mlmodelc, TextDecoder.mlmodelc,
and MelSpectrogram.mlmodelc, not just a parent named Models. Keep the tokenizer cache too:
for this model and package version, it is under Models/models/openai/whisper-tiny.en/ and includes
tokenizer.json and tokenizer_config.json.
In 0.9.0, download: false controls model downloading. Tokenizer loading can still fall back to
the network if local files are missing or invalid. Test a fully provisioned installation with
network access disabled before promising offline operation. The
pinned tokenizer loader
documents this distinction. For bundled assets on iOS, resolve their real bundle URLs, preserve
the directory structure, and verify the resource membership in your app target.
Memory pressure
Issue: App crashes due to memory limitations.
Solution: Use a smaller model. The example uses the SDK’s chunkingStrategy: .vad and limits
concurrentWorkerCount to 1. This reduces concurrent decoding work; the file loader still reads
the audio into memory. Long recordings can therefore require a separate, bounded audio-splitting
workflow. There is no built-in splitAudioIntoChunks function in this example.
Slow transcription
Issue: Transcription takes too long for your use case.
Solution: Measure model loading separately from transcription, then try a smaller model.
Recognition options belong in DecodingOptions, passed to transcribe as decodeOptions.
WhisperKitConfig has no beamSize property in this pinned version. See the
configuration definitions
for supported options.
Streaming transcription
The program above transcribes existing files. Microphone streaming needs a capture lifecycle, microphone authorization, and app-specific start/stop handling. WhisperKit’s upstream streaming implementation is a starting point for that separate integration. The upstream CLI belongs to WhisperKit’s own source package; adding its library dependency does not install that CLI into this example package.
Transloadit's speech transcription capabilities
If you need a cloud-based solution without managing infrastructure, Transloadit offers a powerful speech transcription service as part of our Artificial Intelligence service. Our 🤖 speech/transcribe Robot transcribes speech in audio or video files. Its documentation lists the supported providers, output formats, languages, and provider-specific options, including diarization.
This cloud-based approach can be ideal for processing large files or when on-device processing isn't feasible.
