Edge AI hardware compatibility matrix
Verified memory footprints, power envelopes, and inference latencies across bare-metal CPUs, Raspberry Pi, Apple Silicon, NVIDIA Jetson, and in-browser WebGPU. 100% air-gapped, zero cloud dependencies.
NanoForecast
v05Runs natively on ARM64 Cortex-A76 with zero throttling. ArXiv 2608.14658.
pico-type
v02Tokenizer-free byte architecture. Classifies 62 programming languages and 90 MIME types.
Bharat-Tiny-LLM
v2/v3 MLXZero cloud latency on MacBook Air/Pro. Powered by Brahmi Tokenizer with 33.8% token compression.
Prajna-V2 CRN
CRNEpisodicMemory retrieval table achieving 100% exact factual recall on verifiable facts.
TinyDoc-VLM
256MExtracts invoices, forms, and tables directly on standard CPU servers without GPU billing.
pico-type
v02Sub-6ms CPU execution on commodity Intel/AMD hardware. Perfect for high-throughput routing.
fugusashi
v1.3In-line API proxy router slashing cloud LLM token expenses by 40%–70%.
PolyWhisper
v1.05 Indic languages with dual LoRA experts and 20× hallucination reduction vs Whisper-Small.
Bharat-Tiny-LLM
v3 WebGPURuns 100% locally in the browser tab via WebGPU. Zero telemetry or server data transit.
Frequently answered deployment questions
Can I run an LLM on a Raspberry Pi 4 or 5?+
Yes. For time-series and sensor forecasting, NanoForecast v05 runs at <10ms streaming latency on Raspberry Pi 4/5 utilizing under 12MB of RAM. For natural language text classification, pico-type v02 runs in <12ms with a 200KB ONNX binary. For generative Indic LLMs, quantized 1.7B models can run via llama.cpp on Raspberry Pi 5 (8GB RAM).
Does TinyDoc-VLM require an NVIDIA GPU or cloud connection?+
No. TinyDoc-VLM-256M was specifically trained on SigLIP and SmolLM2 to run on commodity x86 CPUs and Apple Silicon without GPU acceleration. It handles receipt OCR, invoice field extraction, and document QA with under 180ms page latency entirely on-premise.
What is the fastest way to run Indic AI models on Apple Silicon Macs?+
Bharat-Tiny-LLM v2/v3 MLX is pre-compiled for Apple Silicon unified memory using Apple’s MLX framework. It executes at ~45 tokens/second with under 1.2GB memory footprint on MacBook Air and MacBook Pro, with 0 cloud calls and 100% offline privacy.
How does Bharat-Tiny-LLM run inside a web browser without sending data to servers?+
Bharat-Tiny-LLM v3 WebGPU utilizes Transformers.js and ONNX Runtime Web to compile the model weights directly to the client device’s GPU shaders. Once cached, the model runs 100% offline inside the browser sandbox with zero network telemetry.
What hardware is required to deploy Eulogik Edge on-premise?+
Eulogik Edge deploys on standard bare-metal servers or air-gapped workstations (e.g. Dell PowerEdge, HP ProLiant, or Apple Silicon Mac Studio). No specialized $30K H100 GPU clusters are required; our models are designed to operate within 8GB to 64GB of RAM.
Deploy sovereign edge models on your own hardware
Eulogik Edge deploys air-gapped foundation intelligence directly into your data centers and edge nodes. Comply with DPDP Act, UAE laws, and zero cloud lock-in.