[navigation]
TL;DR
- Google ML Kit is the best generic on-device face detection for Android at zero cost: 133 contour points, classification limited to eyes open and smiling, no identity recognition, no liveness.
- Banuba Face AR SDK is the pick when a face has to be tracked in real time with something drawn on it: the patented Face Kernel builds a 3D head model straight from the camera feed, tracks 37 facial morphs, and returns 68 anchor points plus a mesh of up to 3,308 vertices.
- Amazon Rekognition is the pick for 1:N search across collections of millions of faces, from $0.0010 per image on the Group 1 APIs, cloud only.
- 3DiVi is the pick for on-premises biometric identification and deepfake screening: published NIST FRVT 1:1 TAR of 99.70% to 99.76% and a shipping DEEPFAKE_ESTIMATOR block.
- Banuba Face AR SDK does 1:1 biometric matching and active liveness, not 1:N identification against a database. For "who is this out of five million people", use Amazon Rekognition or 3DiVi.
- Google ML Kit does now have a 3D face mesh (468 points), which older comparisons miss, but it is Android-only, in beta, capped at 2 faces, needs faces within about 2 metres, and returns no head orientation.
- Plan capacity against lab figures: Banuba Face AR SDK publishes single-face tracking of 25 FPS on Android Low and 30 FPS on Android High, at a maximum angle of 80 degrees, in fixed lab conditions.
- Pricing shapes differ more than prices do: Google ML Kit is free, Amazon Rekognition bills per image and per liveness check, Banuba Face AR SDK is MAU-based with a 14-day full-access trial, 3DiVi sells annual or perpetual licences with no per-request fees.
Face Detection, Face Tracking and Face Recognition Are Three Different Jobs
Most of the disappointment in this category comes from buying for the wrong job. If you are evaluating an android face detection api, decide which of these three you need first, because the shortlist changes completely with the answer.
Face detection answers "Is there a face, and where is it in this frame?" It returns a bounding box, usually some landmarks, sometimes contours. This is a commodity problem on Android, and it is free: Google ML Kit does it on-device at no cost.
Face tracking answers "Is that the same face, where is it pointing, and what is it doing, thirty times a second?" It needs head pose, expression data, and temporal stability, because something is going to be rendered on top: a filter, makeup, glasses, a beauty effect. Jitter that is invisible in a bounding box becomes obvious the moment you pin a 3D object to a cheekbone. This is where Banuba Face AR SDK sits.
Face recognition and 1:N identification answer "Who is this person?" That means face vectors, a searchable collection, a false-accept rate you can defend to a compliance team, and usually presentation-attack detection. Amazon Rekognition and 3DiVi are built for this. Banuba Face AR SDK is not: it supports 1:1 biometric matching and active liveness, and does not search a database of identities.
One axis cuts across all three: on-device versus cloud. On-device keeps biometric data off third-party servers, works offline, and has no per-request bill, but it is bounded by the phone in the user's hand. Cloud gives unbounded scale and deep metadata, but adds a round trip that rules out real-time AR and moves biometric data off the device, which is a GDPR and CCPA conversation rather than a technical one. We covered what the on-device and cloud trade-off looks like once you measure it separately.
Parameters We Analyzed
Shiny features aside, we dug into the trade-offs that actually break a release cycle. Here are the criteria we used to separate the enterprise-grade tools from the weekend projects:
Performance & Latency
Whether each engine works on the device or waits for a cloud response, whether it holds a frame rate without overheating the phone, and what hardware produced any published figure.
Feature Set
Beyond a bounding box: expression and attribute data, liveness detection, and whether the SDK returns anything you can render against.
Integration Complexity
The path from a fresh project to a stable native Android build. If the SDK is a nightmare to hook into your codebase, the licence savings won't matter.
Developer Experience & Support
Who provides human support, who publishes working sample projects, and who is still shipping releases.
Privacy
Whether the SDK processes locally or ships facial data to the cloud, which sets how much GDPR or CCPA paperwork you'll face.
Pricing & Licensing
Which models offer predictable costs and which hide bills in per-request fine print.
Top 4 Face Detection SDKs for Android: Compared
What each tool is genuinely best at, what it does not do, and the figures each vendor actually publishes.
Banuba's Face Detection API & SDK
Banuba Face Detection API is built for the tracking job rather than the detection job: it exists so something can be rendered on the face, stably, at frame rate, with no server in the loop. Banuba has built face detection and recognition technology for 10 or more years, trains on diverse datasets across skin tones, ages, and lighting conditions, and counts Samsung, Gucci, Depop, Bermuda, Oonzoo, Daily, ME Group, and ADNOC among its customers.
Technical Deep-Dive: The Face Kernel
- Banuba's patented Face Kernel skips the usual two-step pipeline. Traditional SDKs detect 2D landmarks, then run heavy nonlinear equations to lift them into a 3D head pose. Banuba's core technology builds the 3D head model directly from the camera feed instead.
- 37 facial morphs. The engine tracks 37 facial morphs and produces 68 facial anchor points plus a mesh of up to 3,308 vertices. That is fewer published points than ML Kit's 468-point mesh or 3DiVi's 470-point mesh; the argument is the architecture, not the point count, because morph-driven 3D construction returns head angles and expression data that a bare mesh does not.
- Anti-jitter. Banuba's patented anti-jitter system runs tracking multiple times per frame, which is the difference between an AR mask that sits on the skin and one that shivers on it.
- On-device. Processing happens on the device, so no biometric data has to leave it.
- Face analysis. Banuba Face API detects gender and emotions in real time, plus driver tiredness monitoring, heart rate, face segmentation, hand tracking, body segmentation, and a touchless interface. Banuba's pricing page also lists face shape, hair, facial hair, and eyewear detection, biometric match, pupillary distance, lighting detection, seasonal colour analysis, HRV analysis, and 1,000+ premade filters.
- Recognition and liveness. Banuba Face API supports 1:1 biometric matching and active liveness detection, where the user is prompted to blink, move their head, or perform an action. It does not perform 1:N identification against a database.
Published Lab Performance
The figures below come from the Banuba Face AR SDK Technical Specification. They are lab-measured under fixed lab conditions, so treat them as a planning baseline and test in your own environment.
|
Scenario
|
Android Low
|
Android High
|
iOS Mid/High
|
|
Single-face tracking
|
25 FPS
|
30 FPS
|
30 FPS
|
|
Max head angle
|
80 degrees
|
80 degrees
|
80 degrees
|
|
Max distance
|
170 cm
|
180 cm
|
230 cm
|
Multi-face tracking supports a maximum of 5 faces: on Android, 4 faces run at 23 to 28 FPS and 5 at 22 to 27 FPS. With effects applied (face filters, avatars, beautification, makeup without lipstick), Android runs at 25 to 29 FPS and iOS at 30 FPS.
The canonical Banuba Face AR SDK description quotes 60 FPS on mid-range mobile hardware and a -90 to +90 head-angle range, while the lab table quotes 25 to 30 FPS and a maximum angle of 80 degrees. The 60 FPS figure is Banuba's standing marketing description; the lab table is the published technical specification, so plan capacity and device support against the lab figures and benchmark before you commit.
Platforms, Size and Requirements
Banuba Face AR SDK covers Web, Windows, Mac, Android, iOS, Flutter, React Native, and Unity. Requirements: iOS 13+ or Android 8.0+ (API 26+), OpenGL ES 3.0+, a 1280x720 camera recommended at a minimum of 30 FPS, and WebGL 2.0+ for web. The SDK adds around 15 Mb depending on the feature set enabled, materially more than ML Kit unbundled, so if binary size is your binding constraint, ML Kit is the better pick.
Ideal Use Cases
- E-commerce and virtual try-on: Gucci and the Brazilian beauty retailer Oceane use Banuba to let shoppers try on glasses or makeup. Oceane reported a 32% add-to-cart rate after integrating Banuba's Virtual Try-on SDK.
- Social and entertainment apps at scale: the fandom platform b.stage reached 1 million MAUs using Banuba for its live engagement features.
- Live streaming and video calls: real-time beauty and AR effects applied on-device in the camera pipeline.
- Automotive: driver tiredness monitoring to support driving safety.
- Broader AR ecosystem: Banuba Face AR SDK extends into virtual makeup, eyewear, headwear, and jewelry try-on.
Banuba ships Kotlin and Java sample apps on GitHub that are essentially production-ready camera activities, documentation organised around practical recipes, Banuba Studio so designers can handle AR assets without the dev team, and a community forum alongside direct engineering support.
Pricing & Licensing
Banuba pricing is MAU-based and quoted individually, with no published price points, no named tiers, and no per-request charges once you are integrated, so the bill tracks your user base rather than your call volume. A 14-day free trial gives full SDK access, which is the sensible way to run the benchmarks above yourself. The Banuba Face AR SDK pricing guide explains the model.
When Banuba is not the right pick: if you only need to find faces in static photos, if your app never renders anything on the face, if binary size is your hard constraint, or if you need to identify a person against a database of identities. In those cases, use Google ML Kit for the first three and Amazon Rekognition or 3DiVi for the last.

Google ML Kit
Google ML Kit is the best generic on-device face detection for Android at zero cost, and for a large share of apps it is the correct answer with no further evaluation. It runs on-device and offline, and if you already use CameraX or Jetpack, the integration is close to free in engineering time too.
Performance & Features
- Landmarks and contours. ML Kit returns landmarks and, in contour mode, 133 points. Google states contours are detected for only the most prominent face, and that enabling contour detection means only one face is detected, which also makes tracking IDs useless.
- Classification. Two things only: eyes open and smiling, and both work only on frontal faces, with an Euler Y angle between -18 and 18 degrees.
- Tracking IDs. ML Kit assigns tracking IDs across frames, and Google states explicitly that this "isn't a form of face recognition".
- Latency. Google's published figure for face detection in fast mode is about 60 ms on a Pixel 3.
- App size. About 800 KB unbundled, with the model downloaded through Google Play services, or about 6.9 MB bundled. minSdkVersion is 23.
- No identity, no liveness. There is no face recognition and no liveness or presentation-attack detection anywhere in ML Kit, so for a banking or KYC flow it is not enough on its own.
Face Mesh Detection (beta)
ML Kit is often described as having no 3D mesh. It does: ML Kit Face Mesh Detection generates a mesh of 468 3D points. The constraints matter. It is Android-only, in beta with no SLA and no deprecation policy, FACE_MESH mode is capped at 2 faces; it needs faces within about 2 metres of the camera, and it returns no tracking ID, no head orientation, and no expression classification. Latency is about 14 ms on a Pixel 3, app size impact is about 6.4 MB, bundled only.
Roadmap Signal
: face-detection at 16.1.7 and face-mesh-detection at 16.0.0-beta1, both dated 08/07/2024, with Google's 2025 and 2026 ML Kit investment going into the GenAI APIs. Detection is a mature problem, so that is not a reason to avoid ML Kit, but do not expect the mesh to leave beta on a schedule you can plan around.
Ideal Use Cases
- Basic utility: auto-cropping profile pictures, triggering a shutter when someone smiles, counting faces in a frame.
- MVPs and pilots: budget is zero, and the requirement is detection only.
Pricing & Licensing
Google ML Kit is offered at no cost under Google's terms of service, and running unbundled through Google Play services keeps the APK impact small. Skip it if you need stable multi-face tracking with effects rendered on top, head orientation, or any form of biometric security.
Amazon Rekognition
Amazon Rekognition is the best option here for large-scale 1:N face search. You can build collections of millions of faces and search them, analyse up to 100 faces in a single image, and get metadata depth no on-device SDK matches. The trade is structural: it is cloud-only.
Performance & Features
- The tech. You send an image or video to AWS and their servers do the work. That round trip is fine for a selfie check during onboarding and unusable for real-time AR overlays.
- Scale. Collections of millions of faces, up to 100 faces per image, with 1:N search as the headline capability.
- Face Liveness. Analyses a short selfie video and detects printed photos, digital photos, digital video, 3D masks, and bypass attacks including pre-recorded and deepfake video. Available for React web, native iOS, and native Android.
- Privacy and setup. Biometric data leaves the device, which is a compliance workstream rather than a checkbox, and IAM roles, S3 buckets, and Cognito identities are far more setup than dropping a library into an Android project.
Pricing & Licensing
Rekognition image analysis is tiered per image, not a flat rate. Group 1 APIs (IndexFaces, CompareFaces, SearchUsersByImage and similar) cost $0.0010 for the first 1M images, $0.0008 for the next 4M, $0.0006 for the next 30M, and $0.0004 beyond. Group 2 APIs (DetectFaces, DetectLabels, DetectText, RecognizeCelebrities, DetectProtectiveEquipment) share the same first tiers and drop to $0.00025 past roughly 35M images. Running multiple APIs against one image bills as multiple images, which is the line that surprises people.
Face metadata storage costs $0.00001 per face vector per month. Face Liveness costs $0.015 per check for the first 500,000 checks, $0.0125 for the next 2.5M, and $0.010 beyond.
The free tier lasts 12 months from account creation and covers 1,000 images per month in each of Group 1 and Group 2, plus 1,000 face vector and 1,000 user vector objects per month, and is not offered for Image Properties. Since 15 July 2025, new AWS customers also receive up to $200 in Free Tier credits, usable within 12 months.
Choose Amazon Rekognition when you need to search a large face database or want the deepest metadata without taxing the user's hardware. Skip it for real-time AR, for offline use, or when biometric data cannot leave the device.
3DiVi
3DiVi is the strongest option here for on-premises biometric identification with deepfake screening. It covers 1:1 verification and 1:N identification, deploys online or offline, and is actively developed. 3DiVi Inc. is a Delaware corporation based in Covina, California, with public samples on GitHub.
Performance & Features
- Landmarks. The docs describe three landmark sets, not one: fda returns 21 points and is the default; tddfa returns 68 points; mesh returns 470 3D points. The "up to 468 face landmarks" figure that circulates in comparisons comes from 3DiVi marketing copy, not from the docs.
- Accuracy figures. 3DiVi publishes NIST FRVT 1:1 TAR figures of 99.70% on VISA at FAR 1E-6, 99.76% on Mugshot at FAR 1E-5, and 99.75% on Border at FAR 1E-6, and describes itself as top-ranked in NIST FRVT, but publishes no rank, no category, and no test date, and its benchmarks page lists different, older numbers for the same table. Read the TAR figures, not the ranking language.
- Product headline stats. The Face SDK page states 99.5% identification accuracy at FAR 1e-4, 0.14 ms identification in databases of up to 10K identities, and a 97.94% liveness detection rate at BPCER 1%.
- Liveness. Passive (2D) and active (smile, blink, turn), plus 3D with an IR depth sensor. 3DiVi publishes its own APCER at BPCER 1% figures and claims no third-party PAD certification, so there is no iBeta or ISO/IEC 30107-3 result to lean on.
- Deepfake Estimator. Real and shipping: the DEEPFAKE_ESTIMATOR processing block reached version 3 in release 3.30.0.
- Latency, honestly. 3DiVi publishes no Android detection latency. Its detector timings (3 ms to 46 ms per frame by model) are single-core Intel Xeon E5-2683 v4 desktop CPU figures, and its only mobile figures are template extraction on a Snapdragon 845 and Pixel 3, from 150 ms up. The "30 to 50 ms on Android" figure that circulates for 3DiVi is not published by 3DiVi.
Maintenance and Licensing
3DiVi is actively maintained: release 3.31.0 landed on 24 July 2026, adding a Face Tracker processing block, 2d_ensemble_light v4, ssyx detector v2, CUDA 12 and RTX 50-series support, and migrating all Android libraries to 16 KB memory pages for Google Play. Patch 3.30.1 followed on 3 August 2026, so this is a maintained product, not a parked one.
Licensing is quote-only with no published price points: a 14-day trial and a Developer Pack for pilots, annual or perpetual licences, online or offline deployment, and no per-request fees. There is no free tier.
Choose 3DiVi for high-security applications where the priority is identifying a specific person and screening for deepfakes, especially when the deployment has to stay on your own infrastructure. Skip it for lightweight tools, creative AR, or when you do not have time to learn the Processing Block architecture.
Android Face Detection SDKs Comparison Table
|
Dimension
|
Banuba Face AR SDK
|
Google ML Kit
|
Amazon Rekognition
|
3DiVi Face SDK
|
|
Best at
|
Real-time on-device face tracking with AR and beauty effects
|
Generic on-device face detection at zero cost
|
Large-scale 1:N face search
|
On-premises biometric identification and deepfake screening
|
|
Where it runs
|
On-device (Web, Windows, Mac, Android, iOS, Flutter, React Native, Unity)
|
On-device, Android and iOS
|
Cloud only, network round trip required
|
On-premises, online or offline deployment
|
|
Face data returned
|
68 anchor points, 37 morphs, mesh up to 3,308 vertices, head angles, expressions
|
Landmarks, 133 contour points (one face only), 468-point mesh in Android-only beta
|
Bounding boxes, landmarks, and attributes for up to 100 faces per image
|
21, 68, or 470 points depending on the landmark set selected
|
|
Identity and liveness
|
1:1 biometric matching and active liveness, no 1:N identification
|
None at all
|
1:N search plus Face Liveness with deepfake and mask detection
|
1:1 and 1:N, passive, active and 3D liveness, Deepfake Estimator
|
|
Pricing model
|
MAU-based, quoted individually, 14-day full-access trial, no per-request fees
|
No cost
|
Per image from $0.0010, $0.015 per liveness check, 12-month free tier
|
Annual or perpetual licence, quote only, 14-day trial, no per-request fees
|
Links: Banuba Face AR SDK documentation, Google ML Kit face detection for Android, Amazon Rekognition pricing, 3DiVi Face SDK
Summary
The choice follows the job, not a ranking. If you need to find faces on Android and nothing more, Google ML Kit is the best generic on-device option, and it costs nothing, which makes most other considerations moot. If you need to search a large face database or pull deep metadata, Amazon Rekognition has the scale and the per-image economics, at the cost of a round trip and biometric data leaving the device. If you need identification and deepfake screening on infrastructure you control, 3DiVi publishes the strongest verification figures here and the clearest on-premises licensing. Banuba Face AR SDK is the pick for one job: real-time on-device face tracking with AR and beauty effects in consumer Android apps, where head angles, expression data, multi-face support, and jitter-free rendering decide whether the feature ships. Outside that job, one of the other three will serve you better.
