React Native bindings for Whisper and NVIDIA Parakeet ASR through whisper.cpp.
whisper.cpp: High-performance inference of OpenAI's Whisper automatic speech recognition (ASR) model
![]() |
![]() |
|---|---|
| iOS: Tested on iPhone 13 Pro Max | Android: Tested on Pixel 6 |
| (tiny.en, Core ML enabled, release mode + archive) | (tiny.en, armv8.2-a+fp16, release mode) |
npm install whisper.rnwhisper.rn downloads the pre-built ios/rnwhisper.xcframework and android/src/main/jniLibs from the matching GitHub release during postinstall. Existing downloads are reused, and each archive is verified with SHA-256 before extraction. Set RNWHISPER_SKIP_POSTINSTALL=1 to skip the download (e.g. when building from source), or run npx whisper-rn-download-artifacts to fetch it again.
Please re-run npx pod-install again.
By default, whisper.rn will use pre-built rnwhisper.xcframework for iOS. If you want to build from source, please set RNWHISPER_BUILD_FROM_SOURCE to 1 in your Podfile.
If you want to use medium or large model, the Extended Virtual Addressing capability is recommended to enable on iOS project.
Add proguard rule if it's enabled in project (android/app/proguard-rules.pro):
# whisper.rn
-keep class com.rnwhisper.** { *; }
It's recommended to use ndkVersion = "24.0.8215888" (or above) in your root project build configuration for Apple Silicon Macs. Otherwise please follow this trobleshooting issue.
By default, whisper.rn will use pre-built libraries for Android: the whisper.cpp core (librnwhisper*.so, one per CPU-feature variant, including the Hexagon NPU variant) comes from android/src/main/jniLibs, and only the small JNI/JSI wrapper is compiled against your React Native version. If you want to build from source, please set rnwhisperBuildFromSource to true in android/gradle.properties (or pass -PrnwhisperBuildFromSource=true). Building from source compiles whisper.cpp once per variant, so it takes a while; rnwhisperVariants=rnwhisper,rnwhisper_v8fp16_va_2 narrows it down.
On Snapdragon devices with a Hexagon Tensor Processor (SM8450 / 8 Gen 1 and newer), whisper.rn can run the whisper model on the NPU through ggml's Hexagon backend. It is used automatically when useGpu is on (the default) and the device qualifies; check WhisperContext.gpu / reasonNoGPU after initWhisper. The NPU path always uses flash attention. VAD and Parakeet contexts stay on the CPU.
The pre-built libraries include everything the backend needs: the rnwhisper_v8fp16_va_2_hexagon variant, and the DSP-side libraries libggml-htp-*.so in the package's bin/arm64-v8a/, which whisper.rn's Gradle build copies into your app's src/main/assets/rnwhisper-hexagon/ (the library extracts them at runtime and points the DSP loader at them). Just add the FastRPC loader to your app manifest so the library can open it at runtime:
<uses-native-library android:name="libcdsprpc.so" android:required="false" />When building from source, the Hexagon variant is only compiled if the Hexagon SDK is on the build machine (scripts/setup-hexagon-sdk.sh installs it to ~/.hexagon-sdk/6.4.0.2, or set HEXAGON_SDK_ROOT), and the DSP-side libraries have to be built with yarn build:hexagon-htp (Docker, using the ghcr.io/snapdragon-toolchain/arm64-android image). Without the SDK the from-source build is CPU-only. The example app does this in example/android/app/build.gradle (prepareHTP) and its manifest.
You will need to prebuild the project before using it. See Expo guide for more details.
The Tips & Tricks document is a collection of tips and tricks for using whisper.rn.
import { initWhisper } from 'whisper.rn'
const whisperContext = await initWhisper({
filePath: 'file://.../ggml-tiny.en.bin',
})
const sampleFilePath = 'file://.../sample.wav'
const options = { language: 'en' }
const { stop, promise } = whisperContext.transcribe(sampleFilePath, options)
const { result } = await promise
// result: (The inference text result from audio file)ParakeetContext runs NVIDIA's Parakeet TDT 0.6B v3 model through the Parakeet API included in whisper.cpp. The v3 model supports English plus 24 other European languages.
Download a GGUF model from ggml-org/parakeet-GGUF before initializing the context. The example app downloads these models at runtime because they are too large to bundle comfortably:
| Model | Approximate size |
|---|---|
ggml-parakeet-tdt-0.6b-v3-q4_0.bin |
356 MB |
ggml-parakeet-tdt-0.6b-v3-q4_k.bin |
416 MB |
ggml-parakeet-tdt-0.6b-v3-q8_0.bin |
669 MB |
ggml-parakeet-tdt-0.6b-v3-f16.bin |
1.26 GB |
import { initParakeet } from 'whisper.rn'
const parakeetContext = await initParakeet({
filePath: 'file://.../ggml-parakeet-tdt-0.6b-v3-q4_0.bin',
useGpu: true,
})
const { stop, promise } = parakeetContext.transcribe(
'file://.../sample.wav',
{ maxThreads: 4 },
)
const { result, segments, isAborted } = await promise
// Cancel an in-flight transcription when needed:
// await stop()
await parakeetContext.release()Parakeet file and base64 inputs must be WAV containing 16-bit PCM audio. transcribeData() accepts raw signed 16-bit PCM as a base64 string or ArrayBuffer; raw audio must be mono at 16 kHz. Compressed formats such as MP3, AAC, and FLAC are not decoded.
Voice Activity Detection allows you to detect speech segments in audio data using the Silero VAD model.
import { initWhisperVad } from 'whisper.rn'
const vadContext = await initWhisperVad({
filePath: require('./assets/ggml-silero-v6.2.0.bin'), // VAD model file
useGpu: true, // Use GPU acceleration (iOS only, VAD stays on the CPU on Android)
nThreads: 4, // Number of threads for processing
})// Detect speech in audio file (supports same formats as transcribe)
const segments = await vadContext.detectSpeech(require('./assets/audio.wav'), {
threshold: 0.5, // Speech probability threshold (0.0-1.0)
minSpeechDurationMs: 250, // Minimum speech duration in ms
minSilenceDurationMs: 100, // Minimum silence duration in ms
maxSpeechDurationS: 30, // Maximum speech duration in seconds
speechPadMs: 30, // Padding around speech segments in ms
samplesOverlap: 0.1, // Overlap between analysis windows
})
// Also supports:
// - File paths: vadContext.detectSpeech('path/to/audio.wav', options)
// - HTTP URLs: vadContext.detectSpeech('https://example.com/audio.wav', options)
// - Base64 WAV: vadContext.detectSpeech('data:audio/wav;base64,...', options)
// - Assets: vadContext.detectSpeech(require('./assets/audio.wav'), options)// Detect speech in base64 encoded float32 PCM data
const segments = await vadContext.detectSpeechData(base64AudioData, {
threshold: 0.5,
minSpeechDurationMs: 250,
minSilenceDurationMs: 100,
maxSpeechDurationS: 30,
speechPadMs: 30,
samplesOverlap: 0.1,
})segments.forEach((segment, index) => {
console.log(
`Segment ${index + 1}: ${segment.t0.toFixed(2)}s - ${segment.t1.toFixed(
2,
)}s`,
)
console.log(`Duration: ${(segment.t1 - segment.t0).toFixed(2)}s`)
})await vadContext.release()
// Or release all VAD contexts
await releaseAllWhisperVad()The new RealtimeTranscriber provides enhanced realtime transcription with features like Voice Activity Detection (VAD), auto-slicing, and memory management.
// If your RN packager is not enable package exports support, use whisper.rn/src/realtime-transcription
import { RealtimeTranscriber } from 'whisper.rn/realtime-transcription'
import { AudioPcmStreamAdapter } from 'whisper.rn/realtime-transcription/adapters'
import RNFS from 'react-native-fs' // or any compatible filesystem
// Dependencies
const whisperContext = await initWhisper({
/* ... */
})
const vadContext = await initWhisperVad({
/* ... */
})
const audioStream = new AudioPcmStreamAdapter() // requires @fugood/react-native-audio-pcm-stream
// Create transcriber
const transcriber = new RealtimeTranscriber(
{ whisperContext, vadContext, audioStream, fs: RNFS },
{
audioSliceSec: 30,
vadPreset: 'default',
autoSliceOnSpeechEnd: true,
transcribeOptions: { language: 'en' },
},
{
onTranscribe: (event) => console.log('Transcription:', event.data?.result),
onVad: (event) => console.log('VAD:', event.type, event.confidence),
onStatusChange: (isActive) =>
console.log('Status:', isActive ? 'ACTIVE' : 'INACTIVE'),
onError: (error) => console.error('Error:', error),
},
)
// Start/stop transcription
await transcriber.start()
await transcriber.stop()To use Parakeet, provide parakeetContext instead of whisperContext:
const parakeetContext = await initParakeet({
filePath: 'file://.../ggml-parakeet-tdt-0.6b-v3-q4_0.bin',
})
const transcriber = new RealtimeTranscriber(
{ parakeetContext, vadContext, audioStream, fs: RNFS },
{ transcribeOptions: { maxThreads: 4, audioCtx: 0 } },
{ onTranscribe: (event) => console.log(event.data?.result) },
)initialPrompt and promptPreviousSlices are Whisper-only and are ignored when using parakeetContext. Realtime Parakeet audio must be mono, 16 kHz, signed 16-bit PCM.
Long-running sessions: maxSlicesInMemory only bounds raw audio. Per-slice results (transcript + segments) returned by getTranscriptionResults() are kept until stop()/reset() by default, so a transcriber that runs for hours or days grows without bound. Set maxResultsInMemory to keep only the newest N results (older ones are dropped oldest-first) and, when promptPreviousSlices is enabled, maxPromptSlices to cap how many previous results are appended to each Whisper prompt. Consumers that need every result should collect them from onTranscribe as they arrive.
const transcriber = new RealtimeTranscriber(
{ whisperContext, vadContext, audioStream },
{
maxSlicesInMemory: 5, // raw audio slices
maxResultsInMemory: 100, // detailed per-slice results
promptPreviousSlices: true,
maxPromptSlices: 3, // previous results per prompt
},
callbacks,
)Dependencies:
@fugood/react-native-audio-pcm-streamforAudioPcmStreamAdapter- Compatible filesystem module (e.g.,
react-native-fs). See filesystem interface for TypeScript definition
Custom Audio Adapters: You can create custom audio stream adapters by implementing the AudioStreamInterface. This allows integration with different audio sources or custom audio processing pipelines.
Example: See complete example for full implementation including file simulation and UI.
Please visit the Documentation for more details.
You can also use the model file / audio file from assets:
import { initWhisper } from 'whisper.rn'
const whisperContext = await initWhisper({
filePath: require('../assets/ggml-tiny.en.bin'),
})
const { stop, promise } = whisperContext.transcribe(
require('../assets/sample.wav'),
options,
)
// ...This requires editing the metro.config.js to support assets:
// ...
const defaultAssetExts = require('metro-config/src/defaults/defaults').assetExts
module.exports = {
// ...
resolver: {
// ...
assetExts: [
...defaultAssetExts,
'bin', // whisper.rn: ggml model binary
'mil', // whisper.rn: CoreML model asset
],
},
}Please note that:
- It will significantly increase the size of the app in release mode.
- The RN packager is not allowed file size larger than 2GB, so it not able to use original f16
largemodel (2.9GB), you can use quantized models instead.
Platform: iOS 15.0+, tvOS 15.0+
To use Core ML on iOS, you will need to have the Core ML model files.
The .mlmodelc model files is load depend on the ggml model file path. For example, if your ggml model path is ggml-tiny.en.bin, the Core ML model path will be ggml-tiny.en-encoder.mlmodelc. Please note that the ggml model is still needed as decoder or encoder fallback.
The Core ML models are hosted here: https://huggingface.co/ggerganov/whisper.cpp/tree/main
If you want to download model at runtime, during the host file is archive, you will need to unzip the file to get the .mlmodelc directory, you can use library like react-native-zip-archive, or host those individual files to download yourself.
The .mlmodelc is a directory, usually it includes 5 files (3 required):
[
'model.mil',
'coremldata.bin',
'weights/weight.bin',
// Not required:
// 'metadata.json', 'analytics/coremldata.bin',
]Or just use require to bundle that in your app, like the example app does, but this would increase the app size significantly.
const whisperContext = await initWhisper({
filePath: require('../assets/ggml-tiny.en.bin')
coreMLModelAsset:
Platform.OS === 'ios'
? {
filename: 'ggml-tiny.en-encoder.mlmodelc',
assets: [
require('../assets/ggml-tiny.en-encoder.mlmodelc/weights/weight.bin'),
require('../assets/ggml-tiny.en-encoder.mlmodelc/model.mil'),
require('../assets/ggml-tiny.en-encoder.mlmodelc/coremldata.bin'),
],
}
: undefined,
})In real world, we recommended to split the asset imports into another platform specific file (e.g. context-opts.ios.js) to avoid these unused files in the bundle for Android.
The example app provide a simple UI for testing the functions.
Used Whisper model: tiny.en in https://huggingface.co/ggerganov/whisper.cpp
Sample file: jfk.wav in https://github.com/ggerganov/whisper.cpp/tree/master/samples
Please follow the Development Workflow section of contributing guide to run the example app.
We have provided a mock version of whisper.rn for testing purpose you can use on Jest:
jest.mock('whisper.rn', () => require('whisper.rn/jest-mock'))- BRICKS: Our product for building interactive signage in simple way. We provide LLM functions as Generator LLM/Assistant.
- ... (Any Contribution is welcome)
- whisper.node: An another Node.js binding of
whisper.cppbut made API same aswhisper.rn.
See the contributing guide to learn how to contribute to the repository and the development workflow.
See the troubleshooting if you encounter any problem while using whisper.rn.
MIT
Made with create-react-native-library
Built and maintained by BRICKS.

