iOS OCR with Apple Vision and Swift
Use Apple Vision’s RecognizeTextRequest to extract text from an image on an iPhone. This walkthrough
builds a small SwiftUI app that opens an image from Files and displays selectable text. It handles
rotated photos, empty results, unreadable files, and cancellation when the user chooses another
image or leaves the app.
Which iOS OCR API should you use?
For still images, Vision’s Swift API provides asynchronous text recognition on iOS 18 and later. The example below uses this API and English text. No third-party OCR SDK or downloaded model file is needed.
If you support earlier iOS versions, investigate
VNRecognizeTextRequest
instead; the code below is not a compatibility wrapper. For a live camera interface, see
the VisionKit guidance below.
Recognize text in a still image
Create an iOS app project in Xcode named ImageOCR, using SwiftUI and Swift without a storage
integration. Set its deployment target to iOS 18.0 and use Swift 6 language mode. Build with
Xcode 16 or later
for this API. The example was built with Xcode 27.0, the Swift 6.4 compiler in Swift 6 mode, and the iOS
27.0 SDK, then run on an iPhone 17 simulator with iOS 27.0. iOS 18 is the API minimum, not a runtime
tested here.
You will use three files: add OCR.swift, replace ContentView.swift, and replace
ImageOCRApp.swift. Keep all three in the app target. There are no package dependencies or camera
and photo-library permission keys to add because the app selects a file through the system chooser.
Start with a local JPEG or PNG containing large, clear English text. If your image is in Photos,
use its share sheet to save it to Files first.
Read the image and recognize its text
Put this in OCR.swift. The recognizer actor keeps file reading and image decoding off the main
actor. It reads the first image in the file and passes its
Image I/O orientation
to Vision. A CGImage alone contains the pixels, not the orientation needed to display them upright.
import Foundation
import ImageIO
import Observation
import Vision
enum ImageInputError: Error {
case unreadableImage
}
actor TextRecognizer {
func recognize(_ url: URL) async throws -> String {
try Task.checkCancellation()
let hasAccess = url.startAccessingSecurityScopedResource()
defer {
if hasAccess { url.stopAccessingSecurityScopedResource() }
}
let data = try Data(contentsOf: url)
guard let source = CGImageSourceCreateWithData(data as CFData, nil),
let image = CGImageSourceCreateImageAtIndex(source, 0, nil)
else {
throw ImageInputError.unreadableImage
}
let properties = CGImageSourceCopyPropertiesAtIndex(source, 0, nil) as? [CFString: Any]
let rawOrientation = properties?[kCGImagePropertyOrientation] as? UInt32 ?? 1
let orientation = CGImagePropertyOrientation(rawValue: rawOrientation) ?? .up
try Task.checkCancellation()
var request = RecognizeTextRequest()
request.recognitionLevel = .accurate
request.recognitionLanguages = [Locale.Language(identifier: "en-US")]
request.usesLanguageCorrection = true
let observations = try await request.perform(on: image, orientation: orientation)
try Task.checkCancellation()
return observations.compactMap { $0.topCandidates(1).first?.string }
.joined(separator: "\n")
}
}
@MainActor
@Observable
final class OCRModel {
var text = ""
var message = "Choose an image to recognize."
var isRecognizing = false
private let recognizer = TextRecognizer()
private var work: Task<Void, Never>?
func importImage(_ result: Result<URL, Error>) {
cancel()
text = ""
guard case let .success(url) = result else {
message = "Could not open the selected file."
return
}
message = "Recognizing…"
isRecognizing = true
work = Task {
do {
let recognized = try await recognizer.recognize(url)
try Task.checkCancellation()
text = recognized
message = recognized.isEmpty ? "No text found." : "Text recognized."
} catch {
// A canceled task must not replace feedback from a newer selection.
guard !Task.isCancelled else { return }
message = "Could not read this image. Try a local JPEG or PNG."
}
isRecognizing = false
work = nil
}
}
func cancel() {
guard isRecognizing else { return }
work?.cancel()
work = nil
isRecognizing = false
message = "Recognition canceled."
}
}
The model distinguishes a valid image with no recognized text from a file that could not be read. Cancellation is cooperative: it prevents a late result from changing the screen, but does not promise that native processing stops immediately. UI changes stay on the main actor, with no suspension between the final cancellation check and publishing the result.
Connect the Files chooser and results
Replace ContentView.swift with the following. The
fileImporter contract
requires scoped access while reading the selected URL; TextRecognizer balances that access with
defer. A URL already inside the app’s sandbox can still be readable when starting scoped access
returns false, so the file read determines success.
import SwiftUI
import UniformTypeIdentifiers
struct ContentView: View {
@Environment(\.scenePhase) private var scenePhase
@State private var model = OCRModel()
@State private var showImporter = false
var body: some View {
NavigationStack {
VStack(alignment: .leading, spacing: 16) {
HStack {
Button("Choose image") {
model.cancel()
showImporter = true
}
Button("Cancel recognition") { model.cancel() }
.disabled(!model.isRecognizing)
}
Text(model.message)
if model.isRecognizing {
ProgressView()
}
ScrollView {
Text(model.text)
.frame(maxWidth: .infinity, alignment: .leading)
.textSelection(.enabled)
}
}
.padding()
.navigationTitle("Image OCR")
}
.fileImporter(isPresented: $showImporter, allowedContentTypes: [.image]) { result in
model.importImage(result)
}
.onChange(of: scenePhase) { _, phase in
if phase != .active { model.cancel() }
}
.onDisappear { model.cancel() }
}
}
Replace ImageOCRApp.swift with the app entry point:
import SwiftUI
@main
struct ImageOCRApp: App {
var body: some Scene {
WindowGroup { ContentView() }
}
}
Run the app and tap Choose image. Browse to your image in Files and select it. The app clears the previous transcript, shows Recognizing…, then displays Text recognized. and the extracted text. Long-press the transcript to select or copy it. A blank image instead produces No text found.. The app retains the current transcript in memory; it does not save an output file or upload it.
Opening the chooser and selecting a file are separate actions. From an idle screen, opening and
then dismissing the chooser keeps the previous result. During recognition, tapping
Choose image immediately cancels the pending task. Dismissing
that chooser leaves Recognition canceled.; selecting a file
starts new work. The system does not call fileImporter’s completion handler on cancellation.
Cancel recognition also cancels pending work. The app cancels when its scene becomes inactive, including on the way to the background. Returning to the app does not restart recognition: choose the image again. Completed text remains visible when the app returns, provided the process has stayed alive.
Check the result and failure states
Try a simple image containing SWIFT VISION 2468 before testing a complex receipt. That text was
recognized by this example on the simulator, including when its pixels were rotated and the file
carried the matching orientation metadata. OCR can still misread real documents; the newline-joined
observations are not a reconstruction of tables or page layout.
| Input or action | Expected behavior |
|---|---|
| Clear English text | A selectable transcript and Text recognized. |
| Blank image | No text found. with an empty transcript |
| Damaged image or a file that disappears before reading | Could not read this image. Try a local JPEG or PNG. |
| Chooser reports an import error | Could not open the selected file. |
| Cancel or replace pending work | The old task cannot publish a result over the new feedback |
For an unreadable file, first try copying it to a local Files folder and reopening it. A cloud file provider may need connectivity to deliver the bytes even though recognition itself runs locally. This sample reads the whole image into memory; resize unusually large inputs before using it as a batch-processing component.
Scan text from the live camera
DataScannerViewController
is VisionKit’s camera interface for recognizing text and codes. Before offering it, check both
isSupported and isAvailable, provide NSCameraUsageDescription, and handle scanning becoming
unavailable while the app is running. Keep the still-image chooser as a fallback.
A live-camera feature needs testing on supported physical hardware. In particular, deny camera access, grant it in Settings, and return to the existing app process to check recovery. Also test interruption and dismissal while scanning. The simulator exercise above does not establish any of those camera behaviors; use Apple’s linked integration guide when adding that separate feature.
Move batch OCR to a Transloadit workflow
The example above serves one interaction in a running app. If your inputs already belong in an upload-processing workflow, consider Transloadit’s image OCR documentation for the server-side path. That is a separate integration from this local app. Keep service secrets on your backend, outside the iOS bundle.
Improve OCR accuracy
Start with orientation, focus, and readable text size. Cropping away irrelevant surroundings can make a document easier to recognize; stretching a tiny source image cannot restore missing detail. For photographs, reduce glare and steep perspective before adjusting recognition settings.
If a dense page produces no text, check
minimumTextHeightFraction.
Its default is 1/32 of the image height, so small print can be excluded even when it looks clear.
Crop to a smaller region or lower that threshold; considering smaller text can increase recognition
time and memory use.
The example selects English and .accurate. For other languages, inspect the request’s
supportedRecognitionLanguages
and set recognitionLanguages to suit your input. Vision also exposes customWords for vocabulary
such as product names. Test language correction against your own identifiers: a plausible word is
not necessarily the correct serial number.
Common questions
What is the best OCR library for iOS?
Start with Apple Vision when its language coverage fits your images and deployment target. It avoids adding a third-party native OCR binary. If you need a custom model or an unsupported language, evaluate that requirement with representative images before choosing another engine.
Should I use SwiftyTesseract in a new iOS app?
SwiftyTesseract is archived, and its maintainer states that it will receive no further updates. Do not treat it as a maintained dependency for a new app. An existing integration needs its own migration and compatibility testing.
Does iOS OCR require a network connection?
Vision performs text recognition on the device. This app has no model-download or server setup step. Choose a file already stored locally when working offline; downloading a file from iCloud or another provider is a separate operation.
How should I test OCR?
Run both recognition and the surrounding app flow. Include known text, blank and damaged files, rotated pixels with orientation metadata, cancellation, replacement, and background/foreground transitions. Use your actual languages, fonts, and layouts for acceptance tests. Check required fields and useful output, rather than expecting one exact transcript across every OS release.
