I built an Android fatigue reminder app that watches for closed eyes and an open mouth through the front camera and does the detection on the phone itself. Once the model is wired in, the real work is the judgement that comes after it. Alert on a single blink and the reminders quickly turn into noise; keep showing normal when the camera can't see anything and you mislead people about what the system can do.
First, measure the eyes and the mouth
The project uses ML Kit Face Detection, with the model bundled into the APK so there is no download to wait for after installation. It provides the face contour, mouth landmarks and eye-open probabilities, while the fatigue rules are implemented in the app. That keeps detection and judgement separate: changing a duration does not mean touching the camera code.

The eyes use EAR, the eye aspect ratio. Take six points from the eye contour, add the two upper-to-lower eyelid distances, and divide by twice the eye width. The formula is EAR = (one eyelid distance + the other) ÷ (2 × eye width). When the eye closes, the height shrinks and EAR usually drops. Working with a ratio reduces the effect of how far away the face is, but side angles, occlusion and landmark drift still affect the result.
The mouth uses MAR. In this project it is the distance from the centre of the upper outer lip to the centre of the lower outer lip, divided by the distance between the mouth corners. The formula looks simple, but where you put the points is not something you can wave away. Measuring the outer lip and measuring the opening inside the mouth give you different numbers; copying a threshold off the internet is quite likely to be wrong for your case.
I also kept ML Kit's eye-open probability and OR it with EAR: if either path meets the condition, the matching timer starts. The probability path requires valid values for both eyes. That lets both signals take part on their own, but it also means an error in one path can trigger a candidate; they come from the same detector, so they are not two independent pieces of evidence. The mouth is judged separately, so when the eyes are temporarily unusable a valid MAR can still take part.
Crossing a line once does not count as an action yet
A single frame only tells you how narrow the eyes are and how wide the mouth is at that instant. Separating a blink, a brief mouth opening and a sustained action needs time on top of that. The current ordinary eye-closure condition is EAR ≤ 0.22, or either eye below 0.35 when both probabilities are valid, held for 800 ms to enter WARNING. The stronger condition is 0 < EAR < 0.17, or both probabilities below 0.35, held for 400 ms to enter DANGER, which takes priority.
The yawn candidate uses two thresholds. The timer only starts once MAR passes 0.40, and after that it keeps running as long as MAR stays at or above 0.35, reaching 800 ms cumulative before it reminds you. If the value wobbles between 0.41, 0.39 and 0.42, a single threshold would reset the timer over and over. The mouth is still open, but the timer has closed itself for you.
The key branch in the code is only a few lines. Below is an excerpt from the decision logic, with the invalid-value checks and the timing that follows omitted.
val yawning = mar != null && if (highMarStartMs == null) {
mar > config.marYawnEnterThreshold // 0.40
} else {
mar >= config.marYawnExitThreshold // 0.35
}
The time comes from the camera frame's timestamp, not from the moment the model callback finishes. Otherwise a slower inference this time and a faster one next time would be mixed into the action duration. When two valid observations are more than 1.5 seconds apart, the unfinished timer is cleared and starts again from the current frame. With no frames in between, the system has no basis for joining two mouth openings into one continuous action. That interval is an engineering trade-off too; it cannot prove what happened between samples.
If it cannot see clearly, don't keep reporting normal
Besides NORMAL, WARNING and DANGER, the judge keeps a DEGRADED state. When there is no face, when all the eye and mouth signals are invalid, or when a dark frame or a camera error is detected, it clears the unfinished timers and shows the reason it currently cannot detect reliably. NORMAL also only means that the current valid signals have not met the reminder condition; it must not be read as the driver being alert.
On the camera side, acquireLatestImage() takes the newest analysis frame, and if the current session already has an inference task running, the new frame is released immediately so the app does not fall further and further behind. When detection stops or the app goes to the background, the session number increases; an old task that returns late must not overwrite the newer state. After detection restarts, the timing starts again from the new frames as well.
Sound and vibration are only triggered by DANGER; WARNING just updates the interface. The danger alert has a cooldown of about two seconds and is sent through a SharedFlow that does not replay historical events; the interface only receives it while resumed in the foreground, so pausing stops the prompts. Otherwise, coming back to the app would suddenly fire a reminder about something that already happened.
As it stands, this implementation is still a research prototype with fixed thresholds and no personal calibration. A sustained open mouth can also come from talking, and the eye-open probability is limited by the face angle. On the phone side I have only verified installation, the camera pipeline and foreground/background switching on an emulator; real-person detection, overlay accuracy, and the sound and vibration still need testing on a target device. The more worthwhile next step is to record normal speech, brief mouth openings, natural yawns and closed eyes, and use them to check false positives and misses in a static setting; this prototype is not yet a road-validated driving safety system.



