SBC · ARM CortexRaspberry Pi 4 / 5

NanoForecast

v05
Task: Time Series Forecasting
Min RAM
12 MB
Rec. RAM
512 MB
Latency
<10ms streaming
Power
5W–12W (Fanless)

Runs natively on ARM64 Cortex-A76 with zero throttling. ArXiv 2608.14658.

SBC · ARM CortexRaspberry Pi 4 / 5

pico-type

v02
Task: Byte & MIME Classification
Min RAM
8 MB
Rec. RAM
256 MB
Latency
<12ms on ARM
Power
5W

Tokenizer-free byte architecture. Classifies 62 programming languages and 90 MIME types.

Unified Memory · MPSApple Silicon (M1–M4)

Bharat-Tiny-LLM

v2/v3 MLX
Task: Indic Language LLM
Min RAM
1.2 GB
Rec. RAM
8 GB
Latency
45 tok/sec
Power
15W–30W

Zero cloud latency on MacBook Air/Pro. Powered by Brahmi Tokenizer with 33.8% token compression.

Unified Memory · MPSApple Silicon (M1–M4)

Prajna-V2 CRN

CRN
Task: Reasoning · Zero Hallucination
Min RAM
4 GB
Rec. RAM
16 GB
Latency
<35ms recall
Power
20W

EpisodicMemory retrieval table achieving 100% exact factual recall on verifiable facts.

Server / Laptop (No GPU)Commodity x86 CPU

TinyDoc-VLM

256M
Task: Document Vision & OCR
Min RAM
512 MB
Rec. RAM
2 GB
Latency
<180ms page OCR
Power
15W–65W

Extracts invoices, forms, and tables directly on standard CPU servers without GPU billing.

Server / Laptop (No GPU)Commodity x86 CPU

pico-type

v02
Task: Byte & MIME Classification
Min RAM
4 MB
Rec. RAM
128 MB
Latency
<6ms AVX2 CPU
Power
Any CPU

Sub-6ms CPU execution on commodity Intel/AMD hardware. Perfect for high-throughput routing.

Server / Laptop (No GPU)Commodity x86 CPU

fugusashi

v1.3
Task: Federated LLM Routing
Min RAM
256 MB
Rec. RAM
1 GB
Latency
4ms evaluation
Power
Minimal

In-line API proxy router slashing cloud LLM token expenses by 40%–70%.

Edge CUDA · TensorRTNVIDIA Jetson (Orin / Nano)

PolyWhisper

v1.0
Task: Multilingual Indic ASR
Min RAM
1.5 GB
Rec. RAM
4 GB
Latency
<80ms real-time audio
Power
7W–15W

5 Indic languages with dual LoRA experts and 20× hallucination reduction vs Whisper-Small.

WebGPU · 100% OfflineClient-Side In-Browser

Bharat-Tiny-LLM

v3 WebGPU
Task: Offline In-Browser LLM
Min RAM
1.5 GB (Browser Cache)
Rec. RAM
4 GB RAM
Latency
25 tok/sec (WebGPU)
Power
Client battery

Runs 100% locally in the browser tab via WebGPU. Zero telemetry or server data transit.

HARDWARE FAQ

Frequently answered deployment questions

Can I run an LLM on a Raspberry Pi 4 or 5?+

Yes. For time-series and sensor forecasting, NanoForecast v05 runs at <10ms streaming latency on Raspberry Pi 4/5 utilizing under 12MB of RAM. For natural language text classification, pico-type v02 runs in <12ms with a 200KB ONNX binary. For generative Indic LLMs, quantized 1.7B models can run via llama.cpp on Raspberry Pi 5 (8GB RAM).

Does TinyDoc-VLM require an NVIDIA GPU or cloud connection?+

No. TinyDoc-VLM-256M was specifically trained on SigLIP and SmolLM2 to run on commodity x86 CPUs and Apple Silicon without GPU acceleration. It handles receipt OCR, invoice field extraction, and document QA with under 180ms page latency entirely on-premise.

What is the fastest way to run Indic AI models on Apple Silicon Macs?+

Bharat-Tiny-LLM v2/v3 MLX is pre-compiled for Apple Silicon unified memory using Apple’s MLX framework. It executes at ~45 tokens/second with under 1.2GB memory footprint on MacBook Air and MacBook Pro, with 0 cloud calls and 100% offline privacy.

How does Bharat-Tiny-LLM run inside a web browser without sending data to servers?+

Bharat-Tiny-LLM v3 WebGPU utilizes Transformers.js and ONNX Runtime Web to compile the model weights directly to the client device’s GPU shaders. Once cached, the model runs 100% offline inside the browser sandbox with zero network telemetry.

What hardware is required to deploy Eulogik Edge on-premise?+

Eulogik Edge deploys on standard bare-metal servers or air-gapped workstations (e.g. Dell PowerEdge, HP ProLiant, or Apple Silicon Mac Studio). No specialized $30K H100 GPU clusters are required; our models are designed to operate within 8GB to 64GB of RAM.

Deploy sovereign edge models on your own hardware

Eulogik Edge deploys air-gapped foundation intelligence directly into your data centers and edge nodes. Comply with DPDP Act, UAE laws, and zero cloud lock-in.