Implementing OCR in Android apps with Google ML Kit
Use ML Kit’s bundled text recognizer when your Android app needs to read a photo on its first run, including without a network connection. This example builds a small Kotlin app with two inputs: choose an existing image or take a full-size photo with the system camera app. It displays Latin-script text and gives a visible result when the image is blank, unreadable, or canceled.
Prerequisites
You need Kotlin familiarity, an Android SDK installation, and a device or emulator. The example uses Android Gradle plugin 9.2.1, Gradle 9.4.1, JDK 21, SDK Platform 36, and Build Tools 36.0.0. These are a reproducible set of versions, not a requirement to upgrade an existing app. See the AGP compatibility table if you use a different toolchain.
Set JAVA_HOME to your JDK and ANDROID_HOME to your SDK, and put Gradle and the SDK’s
platform-tools directory on your PATH. Building needs network access to fetch dependencies.
The app’s minSdk is 23, following Google’s
Text Recognition v2 setup requirements.
The runtime checks below use Android 16, API 36; they do not establish behavior on every older OS
or physical camera.
Setting up the Android project
Start in a new, empty directory named TextRecognitionApp. Create the files below at the given
relative paths, including their parent directories. These are complete files, so you do not need
to combine them with an Android Studio template or replace files in an existing project.
settings.gradle:
pluginManagement {
repositories {
google()
mavenCentral()
gradlePluginPortal()
}
}
dependencyResolutionManagement {
repositories {
google()
mavenCentral()
}
}
rootProject.name = 'TextRecognitionApp'
include ':app'
build.gradle:
plugins {
id 'com.android.application' version '9.2.1' apply false
}
gradle.properties:
android.useAndroidX=true
org.gradle.jvmargs=-Xmx2g
AGP 9 includes Kotlin compilation, so this project does not apply a separate Kotlin Android plugin.
Adding the ML Kit dependency
Create app/build.gradle. View binding generates ActivityMainBinding from the layout you will
add shortly.
plugins {
id 'com.android.application'
}
android {
namespace 'com.example.textrecognition'
compileSdk 36
defaultConfig {
applicationId 'com.example.textrecognition'
minSdk 23
targetSdk 36
versionCode 1
versionName '1.0'
}
compileOptions {
sourceCompatibility JavaVersion.VERSION_17
targetCompatibility JavaVersion.VERSION_17
}
buildFeatures {
viewBinding true
}
}
dependencies {
implementation 'androidx.activity:activity-ktx:1.9.3'
implementation 'androidx.appcompat:appcompat:1.7.0'
implementation 'com.google.android.gms:play-services-tasks:18.2.0'
implementation 'com.google.mlkit:text-recognition:16.0.1'
}
The last dependency packages the Latin model with the app. The alternative,
com.google.android.gms:play-services-mlkit-text-recognition:19.0.1, downloads its model through
Google Play services. It reduces the initial app download, but recognition cannot return results
until the model is ready. Choose one approach; this walkthrough uses only the bundled model. Google’s
installation comparison
explains the tradeoff and download options.
Configuring permissions
The Photo Picker
grants access to the chosen image without broad photo-library permission. PickVisualMedia uses
ACTION_OPEN_DOCUMENT as its final fallback when a picker is unavailable. Delegating capture to
an installed camera app also avoids requesting direct camera access; Android documents this
camera-intent approach.
Create app/src/main/AndroidManifest.xml:
<manifest xmlns:android="http://schemas.android.com/apk/res/android">
<application
android:label="Text Recognition"
android:theme="@style/Theme.AppCompat.Light.NoActionBar">
<activity android:name=".MainActivity" android:exported="true">
<intent-filter>
<action android:name="android.intent.action.MAIN" />
<category android:name="android.intent.category.LAUNCHER" />
</intent-filter>
</activity>
<provider
android:name="androidx.core.content.FileProvider"
android:authorities="${applicationId}.fileprovider"
android:exported="false"
android:grantUriPermissions="true">
<meta-data
android:name="android.support.FILE_PROVIDER_PATHS"
android:resource="@xml/file_paths" />
</provider>
</application>
</manifest>
Create app/src/main/res/xml/file_paths.xml:
<paths xmlns:android="http://schemas.android.com/apk/res/android">
<cache-path name="ocr_camera" path="ocr-camera/" />
</paths>
FileProvider exposes
only this capture directory through content URIs. TakePicture supplies a URI to the camera so it
can write a full-size image, rather than returning a small thumbnail.
Creating the layout
Create app/src/main/res/layout/activity_main.xml. The result is selectable so you can copy it;
status messages use the same view and remain visible after a failed operation.
<LinearLayout xmlns:android="http://schemas.android.com/apk/res/android"
android:layout_width="match_parent"
android:layout_height="match_parent"
android:orientation="vertical"
android:padding="16dp">
<Button
android:id="@+id/btnCapture"
android:layout_width="match_parent"
android:layout_height="wrap_content"
android:text="Capture image" />
<Button
android:id="@+id/btnGallery"
android:layout_width="match_parent"
android:layout_height="wrap_content"
android:text="Choose image" />
<ScrollView
android:layout_width="match_parent"
android:layout_height="0dp"
android:layout_marginTop="16dp"
android:layout_weight="1">
<TextView
android:id="@+id/textView"
android:layout_width="match_parent"
android:layout_height="wrap_content"
android:text="Choose or capture an image."
android:textIsSelectable="true"
android:textSize="16sp" />
</ScrollView>
</LinearLayout>
Implementing OCR functionality
Create app/src/main/java/com/example/textrecognition/MainActivity.kt. Both buttons stay disabled
while an external activity or recognition is pending. Decoding runs on a worker thread, and ML Kit
receives the selected URI through InputImage.fromFilePath().
Android can recreate your activity while the picker or camera is open. Registering launchers in a stable order and saving the extra operation state lets their results reach the replacement activity. A scan already being recognized is not resumed after recreation: the new screen asks you to select an image again.
package com.example.textrecognition
import android.content.ActivityNotFoundException
import android.net.Uri
import android.os.Bundle
import androidx.activity.enableEdgeToEdge
import androidx.activity.result.PickVisualMediaRequest
import androidx.activity.result.contract.ActivityResultContracts
import androidx.appcompat.app.AppCompatActivity
import androidx.core.content.FileProvider
import androidx.core.view.ViewCompat
import androidx.core.view.WindowInsetsCompat
import com.example.textrecognition.databinding.ActivityMainBinding
import com.google.android.gms.tasks.TaskCompletionSource
import com.google.mlkit.vision.common.InputImage
import com.google.mlkit.vision.text.TextRecognition
import com.google.mlkit.vision.text.latin.TextRecognizerOptions
import java.io.File
import java.io.IOException
import java.util.concurrent.Executors
class MainActivity : AppCompatActivity() {
private lateinit var binding: ActivityMainBinding
private var pendingCameraName: String? = null
private var pendingPicker = false
private var recognizing = false
private val cameraDirectory: File
get() = File(cacheDir, "ocr-camera")
private val takePictureLauncher = registerForActivityResult(
ActivityResultContracts.TakePicture()
) { success ->
val name = pendingCameraName
pendingCameraName = null
if (name == null) {
setBusy(false)
showMessage("The camera result is no longer available.")
} else {
val file = File(cameraDirectory, name)
if (success && file.isFile && file.length() > 0) {
processImage(cameraUri(file), file)
} else {
file.delete()
setBusy(false)
showMessage(if (success) "The camera returned no image." else "Capture canceled.")
}
}
}
private val selectPictureLauncher = registerForActivityResult(
ActivityResultContracts.PickVisualMedia()
) { uri ->
pendingPicker = false
if (uri != null) {
processImage(uri)
} else {
setBusy(false)
showMessage("Selection canceled.")
}
}
override fun onCreate(savedInstanceState: Bundle?) {
super.onCreate(savedInstanceState)
enableEdgeToEdge()
binding = ActivityMainBinding.inflate(layoutInflater)
setContentView(binding.root)
val padding = binding.root.paddingLeft
ViewCompat.setOnApplyWindowInsetsListener(binding.root) { view, insets ->
val bars = insets.getInsets(
WindowInsetsCompat.Type.systemBars() or WindowInsetsCompat.Type.displayCutout()
)
view.setPadding(padding + bars.left, padding + bars.top,
padding + bars.right, padding + bars.bottom)
insets
}
pendingCameraName = savedInstanceState?.getString("pendingCameraName")
pendingPicker = savedInstanceState?.getBoolean("pendingPicker") ?: false
savedInstanceState?.getString("recognizedText")?.let { binding.textView.text = it }
if (savedInstanceState?.getBoolean("recognizing") == true) {
showMessage("Recognition interrupted. Select an image again.")
}
setBusy(pendingCameraName != null || pendingPicker)
// A killed process cannot run its completion listener to delete abandoned captures.
cameraDirectory.listFiles()?.filter {
it.name != pendingCameraName && it.lastModified() < System.currentTimeMillis() - 86_400_000
}?.forEach { it.delete() }
binding.btnCapture.setOnClickListener { captureImage() }
binding.btnGallery.setOnClickListener { selectImage() }
}
private fun selectImage() {
pendingPicker = true
setBusy(true)
showMessage("Waiting for an image…")
try {
selectPictureLauncher.launch(
PickVisualMediaRequest(ActivityResultContracts.PickVisualMedia.ImageOnly)
)
} catch (error: ActivityNotFoundException) {
cancelSelection()
} catch (error: SecurityException) {
cancelSelection()
}
}
private fun cancelSelection() {
pendingPicker = false
setBusy(false)
showMessage("The image picker could not be opened.")
}
private fun captureImage() {
try {
if (!cameraDirectory.isDirectory && !cameraDirectory.mkdirs()) {
throw IOException("Could not create capture directory")
}
val imageFile = File.createTempFile("IMG_", ".jpg", cameraDirectory)
pendingCameraName = imageFile.name
setBusy(true)
showMessage("Waiting for the camera…")
takePictureLauncher.launch(cameraUri(imageFile))
} catch (error: IOException) {
cancelCapture()
} catch (error: ActivityNotFoundException) {
cancelCapture()
} catch (error: SecurityException) {
cancelCapture()
}
}
private fun cameraUri(file: File): Uri =
FileProvider.getUriForFile(this, "${packageName}.fileprovider", file)
private fun cancelCapture() {
pendingCameraName?.let { File(cameraDirectory, it).delete() }
pendingCameraName = null
setBusy(false)
showMessage("The camera could not be opened.")
}
private fun processImage(uri: Uri, temporaryFile: File? = null) {
recognizing = true
setBusy(true)
showMessage("Recognizing text…")
val recognizer = TextRecognition.getClient(TextRecognizerOptions.DEFAULT_OPTIONS)
val decoded = TaskCompletionSource<InputImage>()
val executor = Executors.newSingleThreadExecutor()
executor.execute {
try {
decoded.setResult(InputImage.fromFilePath(applicationContext, uri))
} catch (error: Exception) {
decoded.setException(error)
}
}
executor.shutdown()
decoded.task.continueWithTask { task -> recognizer.process(task.result) }
.addOnSuccessListener { visionText ->
if (!isDestroyed) {
showMessage(visionText.text.ifBlank { "No text found. Try a clearer image." })
}
}
.addOnFailureListener {
if (!isDestroyed) showMessage("The image could not be read. Try another image.")
}
.addOnCompleteListener {
// Activity-scoped listeners stop on onStop; this cleanup must still run.
recognizer.close()
temporaryFile?.delete()
recognizing = false
if (!isDestroyed) setBusy(false)
}
}
private fun setBusy(value: Boolean) {
binding.btnCapture.isEnabled = !value
binding.btnGallery.isEnabled = !value
}
private fun showMessage(message: String) {
binding.textView.text = message
}
override fun onSaveInstanceState(outState: Bundle) {
outState.putString("pendingCameraName", pendingCameraName)
outState.putBoolean("pendingPicker", pendingPicker)
outState.putBoolean("recognizing", recognizing)
outState.putString("recognizedText", binding.textView.text.toString())
super.onSaveInstanceState(outState)
}
}
The inset listener keeps the controls clear of
system bars and display cutouts.
The app keeps completed text across recreation, but does not save a scan history or retain a
preview. Each capture gets a new cache filename. Completed and canceled captures are deleted;
an abandoned capture is removed on a later launch after one day. Leave a pending capture alone in
onDestroy(), because the camera may still be writing it.
The picker image is read for the immediate operation; its original file is never deleted. A recognition task owns its recognizer until completion, even if its activity is destroyed. Its callbacks cannot replace text in a new activity instance.
Optimizing OCR accuracy and performance
Start with a sharp, well-lit photo of printed text. Google’s input image guidelines recommend at least 16 × 16 pixels per character; beyond about 24 × 24, larger characters generally do not improve accuracy. Crop unnecessary background while keeping the text legible.
This is a still-image flow, not a live camera analyzer. A CameraX preview needs its own permission, frame rotation, and backpressure handling. Chinese, Devanagari, Japanese, and Korean also need their corresponding model dependencies and recognizer options rather than the Latin options used here.
Testing the application
From the project directory, build the debug APK. A rerun replaces build outputs, including the APK; it does not overwrite source files.
gradle --no-daemon --max-workers=2 :app:assembleDebug
With your target device connected, replace YOUR_DEVICE_SERIAL with its serial from adb devices.
Installing with -r replaces this sample app while retaining its app data.
adb -s YOUR_DEVICE_SERIAL install -r app/build/outputs/apk/debug/app-debug.apk &&
adb -s YOUR_DEVICE_SERIAL shell am start -n com.example.textrecognition/.MainActivity
Choose a local image containing large text such as “INVOICE 12345”. That text should appear in the scrollable result, and both buttons should become available again. For an offline first-run check, install the app without launching it, disconnect the device’s network, then open it and select a local photo. Cloud-backed picker items may still need a connection to retrieve their bytes.
Check these outcomes before adapting the example:
| Input or interruption | Expected result |
|---|---|
| Blank image | “No text found. Try a clearer image.” |
| Unreadable or missing image | “The image could not be read. Try another image.” |
| Close the picker without selecting | “Selection canceled.”; both buttons enabled |
| Cancel capture | “Capture canceled.”; empty capture file removed |
| Camera reports success without writing bytes | “The camera returned no image.”; capture file removed |
| Rotate while the picker or camera is open | The pending result still belongs to that operation |
| Recreate the activity during recognition | “Recognition interrupted. Select an image again.”; controls enabled |
| Rotate after recognition | Completed text survives |
The example was exercised on an Android 16 emulator with the bundled Latin model. Use real phones for focus, exposure, and camera-app compatibility testing. A rotated JPEG also needs correct orientation metadata; do not assume rotating the phone can repair an incorrectly encoded image.
Troubleshooting
If the app reports no text, first try a clearer image with larger printed characters. A blank recognition result is different from an unreadable file. If capture returns no image, check the camera app and the FileProvider authority, cache path, and manifest resource name together.
A recognizer that only works after going online may be using the unbundled dependency. Check your
resolved Gradle dependencies before adding model-download logic to this bundled example. If the
build cannot resolve ActivityMainBinding, verify the layout filename and view-binding setting.
For an alternative that processes uploaded files on a server, explore the /document/ocr documentation.
