NIRNAY 450M vs Jev API: Banking Intent & Decision Calibration
77-Way Intent Classification, Probability Calibration (ECE), and Edge Inference Economics. Direct empirical comparison between Eulogik’s NIRNAY 450M open-weight calibrated decision model and Jev 1.13.0 API on the 3,080-case Banking77 test benchmark.
EXECUTIVE SUMMARY & EMPIRICAL VERDICT
NIRNAY 450M outperforms Jev API by +7.6 points on Banking77 with calibrated probability outputs, zero data leakage, and zero recurring cloud costs, making it the ideal choice for regulated financial institutions.
EULOGIK EDGE MODELFeatured
NIRNAY
450M (v0.1.0) · 450 Million
Target HardwareApple Silicon M4 / Commodity CPU
Latency / Speed209ms (M4 GPU) · 361ms (CPU)
RAM / Memory466 MB checkpoint
Fertility / SizeOpen-Weight (Apache 2.0)
- ✓ 87.92% accuracy on Banking77 test split (3,080 test cases)
- ✓ Calibrated probabilities: Brier 0.2083, fitted ECE 0.0454
- ✓ Runs 100% locally on your hardware — zero per-call token fees
- ✓ Choice, score, and binary decisions in a single forward pass
- ✓ Zero data leaves your perimeter — air-gapped DPDP & RBI compliant
INDUSTRY BASELINEOpen Weights
Jev API
1.13.0 (Zero-shot) · Proprietary Cloud Model
Target HardwareCloud API / US Data Centers
Latency / Speed310–478ms API network roundtrip
RAM / MemoryCloud Hosted
Fertility / SizeProprietary Cloud API
- • 80.3% accuracy on identical Banking77 test cases
- • Broad general language capability across long documents (>512 tokens)
- • Zero local setup required — accessible via API key
- • Ongoing per-query API billing
- • Sensitive financial customer queries transit external cloud servers
METRIC SPECIFICATION MATRIX
| METRIC | NIRNAY | JEV API |
|---|---|---|
| Banking77 Accuracy | 87.92% ★ | 80.30% |
| Accuracy Delta | +7.62 points ★ | Baseline |
| Probability Calibration (Fitted ECE) | 0.0454 ★ | Not Published |
| Brier Score (lower is better) | 0.2083 ★ | Not Published |
| Data Sovereignty & Privacy | 100% On-Premise / Air-Gapped ★ | Cloud API Transit |
| Per-Call API Cost | $0.00 (Self-Hosted) ★ | Metered API Billing |
| Inference Latency | 209ms (M4 MPS) ★ | 310–478ms (API) |
| Context Window | 512 tokens | 4,000+ tokens ★ |
REPRODUCIBLE CODE SNIPPET
Execute this benchmark locally
from nirnay.agent import NirnayAgent
# Run on-device calibrated decision model
agent = NirnayAgent(device="mps")
out = agent.system_one(
"My card was charged twice for the same transaction at Starbucks.",
questions={
"intent": {
"type": "choice",
"instructions": "Classify banking intent",
"criteria": {"duplicate_charge": "charged twice", "refund_status": "refund"}
}
}
)
print("Confidence:", out["answers"]["intent"]["probabilities"])Interested in deploying NIRNAY?
We assist enterprises with custom fine-tuning, hardware sizing, and on-premise air-gapped installation.