Small models.
Real hardware.
Zero cloud.

355
Maternal deaths per 100K births in Kenya
INT4
Target precision
2B
Parameter ceiling
ARCHITECTURE

Everything engineered
for the edge.

Liquid Foundation Models

LFM2.5 — a hybrid SSM architecture with linear O(N) memory complexity. No expanding KV cache. Constant RAM footprint regardless of context length. Built for edge silicon from the architecture up.

QAT from day one

Fake-quantization nodes integrated into fine-tuning simulate INT4/INT8 constraints during training. The model learns to be accurate under hardware limits — not surprised by them at deployment.

Offline by design

No cloud call. No network dependency. One inference pass, on-device, in Swahili or English. Runs where connectivity cannot be assumed — because that is the design requirement, not an afterthought.

llama.cpp / GGUF export

Single-file export via GGUF. Runs on CPU across Android smartphones, NVIDIA Jetson Orin Nano, and Linux edge boards. No Python runtime required. Maximum portability on minimum hardware.

DEPLOYMENTS

Validated in the
hardest conditions first.

Two proof points. Same methodology. Different domains. Both grounded in real African operating conditions where cloud connectivity cannot be assumed.

MAMA AFYA
MAMA AFYA

Maternal health triage

Offline symptom-to-urgency triage assistant for Community Health Workers in Kenya. Classifies danger signs in Swahili or English entirely on-device — no internet required. Evaluation tracks error severity migration, not just aggregate accuracy.

355
maternal deaths per 100K
INT4
target precision
ZEROBIONIC
ZEROBIONIC

KSL sign language · robotic arm

Fine-tuning LFM2.5-Audio on Swahili/English STEM classroom vocabulary to output Kenyan Sign Language gesture labels executed by Zerobionic's robotic arm — offline, in real time, for deaf students in Kenyan classrooms.

100+
schools reached by Zerobionic
92%
sign language accuracy
COMING NEXT
COMING NEXT

Industrial robotics perception

Applying the same QAT-from-day-one methodology to vision-language perception tasks for robotic systems operating in African industrial environments — mining, ports, agriculture.

sub-2B
parameters
O(N)
memory complexity
OPEN
OPEN

Your deployment

The methodology generalizes. If you are building offline-first assistive technology, healthcare tools, or robotic systems for low-connectivity environments, we want to hear from you.

0
cloud calls
possible deployments
RESEARCH

From hypothesis to
field-validated result.

Hypothesize
01

Hypothesize

QAT-from-day-one preserves more accuracy at INT4 than PTQ. We state this as a falsifiable claim, not an assumption.

Build the dataset
02

Build the dataset

Record locally-grounded Swahili/English audio with Kenyan CHWs and teachers. Align labels to Kenya MoH clinical guidelines and Zerobionic's KSL gesture library.

Benchmark
03

Benchmark

Fine-tune LFM2.5 with QAT vs PTQ baseline. Measure aggregate accuracy, severity-weighted error cost, and catastrophic-drop rate on real edge hardware.

Validate in the field
04

Validate in the field

Mama Afya into 15 CHWs' hands. Zerobionic collaboration into active Kenyan classrooms. Real users, real conditions, results published regardless of outcome.

METHODOLOGY

Beyond aggregate
accuracy.

Standard accuracy metrics mask a critical failure mode. A model with fewer total errors can still concentrate remaining errors on the highest-consequence cases. Vilya's evaluation framework catches this.

Agent orchestration architecture
SEVERITY-WEIGHTED EVAL

Misclassification cost tracking

Misclassifying a high-urgency case as low-urgency carries 100× the cost of an over-referral. We report total weighted error cost alongside aggregate accuracy for every model condition.

Correct → correct
Cost: 0
Low → high urgency
Cost: 1
High → low urgency
Cost: 100
ERROR MIGRATION

For every case correct at FP16, we track whether quantization produces a graceful degradation or a catastrophic drop to the lowest urgency tier.

OPEN QUESTIONS

What we are building
to find out.

Real research requires falsifiable claims, rigorous measurement, and commitment to transparency — regardless of the outcome.

Does QAT survive fine-tuning?

LFM2.5's detokenizer was QAT-optimized before release. Does task-specific fine-tuning reintroduce quantization error that requires a second QAT round to remove?

Does accuracy advantage survive at the case level?

Does QAT produce a lower catastrophic-drop rate on highest-urgency triage cases, or does the advantage exist only in aggregate?

Does tail latency hold?

Does 99th-percentile inference latency remain within an acceptable bound for real-time KSL interpretation on the Jetson Orin Nano?

IN DEVELOPMENT
PRE-DEPLOYMENT
RESULTS PUBLISHED REGARDLESS OF OUTCOME
Development Log
12:34:21QAT fine-tune started
12:34:18PTQ baseline computed
12:34:15Severity matrix applied
12:34:12Error migration tracked
12:34:09On-device benchmark run
Technical Stack

Built on open foundations.

terminal
# Pull LFM2.5-Audio from Hugging Face
$ huggingface-cli download LiquidAI/LFM2.5-Audio-1.5B
Downloading model...
1.5GB / 1.5GB ✓
✓ Model ready
✓ Ready for training
Swahili ASR
KSL Gesture Labels
Maternal Triage
INT4 Inference
Offline Deployment
QAT Fine-tuning
GGUF Export
Edge NPU
Jetson Orin Nano
Android CPU
Swahili ASR
KSL Gesture Labels
Maternal Triage
INT4 Inference
Offline Deployment
QAT Fine-tuning
GGUF Export
Edge NPU
Jetson Orin Nano
Android CPU
Swahili ASR
KSL Gesture Labels
Maternal Triage
INT4 Inference
Offline Deployment
QAT Fine-tuning
GGUF Export
Edge NPU
Jetson Orin Nano
Android CPU
Severity-Weighted Eval
Error Migration
LFM2.5
llama.cpp
Kenya MoH Guidelines
Zerobionic
CHW Tools
Sub-2B Models
Local Data
Zero Cloud
Severity-Weighted Eval
Error Migration
LFM2.5
llama.cpp
Kenya MoH Guidelines
Zerobionic
CHW Tools
Sub-2B Models
Local Data
Zero Cloud
Severity-Weighted Eval
Error Migration
LFM2.5
llama.cpp
Kenya MoH Guidelines
Zerobionic
CHW Tools
Sub-2B Models
Local Data
Zero Cloud
IN THE FIELD

Where Vilya models
will run.

Real hardware. Real users. Real conditions. Real commitment to publishing results regardless of outcome.

2deployments · in active development
AGENTTASKREGIONSTATUS
analyst-7f2a
#A1B2C3
Generating Q2 financial report
us-east
running
executor-3b1c
#D4E5F6
Running integration test suite
eu-west
running
researcher-2c8f
#G7H8I9
Scraping competitor pricing data
us-west
queued
planner-5a3d
#J0K1L2
Syncing Notion docs with Linear
eu-central
running
coder-8d1a
#M3N4O5
Refactoring auth module — 3 files
ap-south
running
monitor-9d4e
#P6Q7R8
Monitoring uptime across 8 regions
us-east
complete
CONTACT

Build with us.

Researchers
  • Working on quantization or small model fine-tuning
  • Edge inference optimization
  • African language NLP
  • Safety-critical AI evaluation
Reach out
Collaborators
  • Deploying models on Jetson or Android NPUs
  • Building assistive or healthcare tools
  • Working in low-connectivity African settings
  • Robotics teams needing an offline inference layer
Let's talk
Partners
  • Organizations funding African AI infrastructure
  • Clinical or health system partners
  • Hardware sponsors (Jetson kits, edge boards)
  • Fellowship and grant programs
Get in touch

Small models.
Real hardware.
Zero cloud.

Vilya Labs builds Small Multimodal Foundation Models engineered with Quantization-Aware Training from day one — running accurately offline on the devices and in the communities that need AI most.