ABOUT THIS MODEL

A 149M-parameter cross-encoder reranker fine-tuned on ModernBERT-base for high-precision retrieval-augmented generation. Evaluates query-document relevance in under 50ms on CPU with an 8K context window. Engineered to separate hard negatives and eliminate hallucinations in legal and financial RAG pipelines without requiring a dedicated GPU.

KEY HIGHLIGHTS
  • ✓Sub-50ms CPU inference for real-time search and retrieval
  • ✓ModernBERT backbone with full 8,192 token context window
  • ✓Fine-tuned on MS MARCO hard negatives for precision reranking
  • ✓Powers zero-GPU on-premise RAG in LexRAG and Evolucent AI
SPECIFICATIONS
Parameters149M
BaseModernBERT-base
Latency<50ms on CPU
Context8,192 tokens
ArchitectureCross-Encoder
HardwareCPU / Edge
LicenseApache 2.0
Downloads21+
QUICK START

Get started in minutes

Install
pip install transformers torch
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch

tokenizer = AutoTokenizer.from_pretrained("eulogik/flashrank-pro-base")
model = AutoModelForSequenceClassification.from_pretrained("eulogik/flashrank-pro-base")

query = "What are the penalties under DPDP Section 16?"
doc = "Section 16 outlines penalties up to 250 crore for significant data fiduciaries..."

inputs = tokenizer(query, doc, return_tensors="pt", truncation=True, max_length=512)
with torch.no_grad():
    scores = model(**inputs).logits
print("Relevance Score:", scores.item())
ACADEMIC & RESEARCH CITATION

Cite this Model & Architecture

If you utilize FlashRank Pro in your academic research, benchmarks, or enterprise deployments, please cite our open-weight publication:

@misc{eulogik2026flashrankpro,
  title        = {FlashRank Pro Base: 149M ModernBERT Cross-Encoder Reranker for Enterprise RAG},
  author       = {Kishore, Gautam and Eulogik Systems Engineering},
  year         = {2026},
  publisher    = {Eulogik},
  howpublished = {\url{https://eulogik.com/models/flashrank-pro}}
}

Need help deploying FlashRank Pro?

We offer consulting, custom fine-tuning, and on-premise deployment for all our models.