LFM2.5 — a hybrid SSM architecture with linear O(N) memory complexity. No expanding KV cache. Constant RAM footprint regardless of context length. Built for edge silicon from the architecture up.
Fake-quantization nodes integrated into fine-tuning simulate INT4/INT8 constraints during training. The model learns to be accurate under hardware limits — not surprised by them at deployment.
No cloud call. No network dependency. One inference pass, on-device, in Swahili or English. Runs where connectivity cannot be assumed — because that is the design requirement, not an afterthought.
Single-file export via GGUF. Runs on CPU across Android smartphones, NVIDIA Jetson Orin Nano, and Linux edge boards. No Python runtime required. Maximum portability on minimum hardware.
Two proof points. Same methodology. Different domains. Both grounded in real African operating conditions where cloud connectivity cannot be assumed.


Offline symptom-to-urgency triage assistant for Community Health Workers in Kenya. Classifies danger signs in Swahili or English entirely on-device — no internet required. Evaluation tracks error severity migration, not just aggregate accuracy.


Fine-tuning LFM2.5-Audio on Swahili/English STEM classroom vocabulary to output Kenyan Sign Language gesture labels executed by Zerobionic's robotic arm — offline, in real time, for deaf students in Kenyan classrooms.


Applying the same QAT-from-day-one methodology to vision-language perception tasks for robotic systems operating in African industrial environments — mining, ports, agriculture.


The methodology generalizes. If you are building offline-first assistive technology, healthcare tools, or robotic systems for low-connectivity environments, we want to hear from you.

QAT-from-day-one preserves more accuracy at INT4 than PTQ. We state this as a falsifiable claim, not an assumption.

Record locally-grounded Swahili/English audio with Kenyan CHWs and teachers. Align labels to Kenya MoH clinical guidelines and Zerobionic's KSL gesture library.

Fine-tune LFM2.5 with QAT vs PTQ baseline. Measure aggregate accuracy, severity-weighted error cost, and catastrophic-drop rate on real edge hardware.

Mama Afya into 15 CHWs' hands. Zerobionic collaboration into active Kenyan classrooms. Real users, real conditions, results published regardless of outcome.
Standard accuracy metrics mask a critical failure mode. A model with fewer total errors can still concentrate remaining errors on the highest-consequence cases. Vilya's evaluation framework catches this.

Misclassifying a high-urgency case as low-urgency carries 100× the cost of an over-referral. We report total weighted error cost alongside aggregate accuracy for every model condition.
For every case correct at FP16, we track whether quantization produces a graceful degradation or a catastrophic drop to the lowest urgency tier.
Real research requires falsifiable claims, rigorous measurement, and commitment to transparency — regardless of the outcome.
LFM2.5's detokenizer was QAT-optimized before release. Does task-specific fine-tuning reintroduce quantization error that requires a second QAT round to remove?
Does QAT produce a lower catastrophic-drop rate on highest-urgency triage cases, or does the advantage exist only in aggregate?
Does 99th-percentile inference latency remain within an acceptable bound for real-time KSL interpretation on the Jetson Orin Nano?
Real hardware. Real users. Real conditions. Real commitment to publishing results regardless of outcome.
Vilya Labs builds Small Multimodal Foundation Models engineered with Quantization-Aware Training from day one — running accurately offline on the devices and in the communities that need AI most.