M37Labs Logo
M37Labs Saransh 1.7B Sovereign Summarization Model
M37LABS SLM · 1.72B COMPACT INTELLIGENCE · APACHE 2.0

Summarization that follows your length.

Generalist LLMs ignore word count constraints and leak sensitive corporate data to external clouds. Saransh 1.7B is an edge-optimized sovereign model built for strict length budgets and calibrated synthesis on local hardware.

01/Model Scale
1.72BActive Parameters
02/Context Window
8,192Trained Tokens
03/Edge Footprint
1.1 GB4-bit Q4_K_M GGUF
04/Short Adherence
100%23.9w avg vs 252w base
05/ROUGE-1 Gain
+104%0.387 vs 0.190 base
06/Open License
Apache 2.0Commercial permissive
THE ARCHITECTURAL DILEMMA

Small model. Strict word budgets. Zero cloud egress.

Saransh (सारांश, Sanskrit for “summary”) was engineered to solve the primary failure mode of small models: uncontrolled verbosity. Untuned base models spew 250+ rambling words when asked for brief summaries, hallucinate phantom sources, and exhaust inference budgets.

01 / INGESTION

One to Three Pages Uncut

Paste dense financial reports, earnings filings, or technical transcripts. With an 8,192 trained token window, documents are ingested as cohesive narratives without breaking paragraphs into fragmented RAG chunks.

8K Trained Context · 40K Arch Maximum
02 / CALIBRATION

Strict Length Adherence

Request a two-sentence headline and get exactly two sentences. Choose from four discrete length regimes or provide exact word count budgets to fit strict UI layout constraints without truncation artifacts.

Short (24w) · Medium (87w) · Long (347w) · Explicit (~120w)
03 / SOVEREIGNTY

Zero Cloud Egress

Quantized down to 1.1 GB in 4-bit GGUF. Runs natively on standard laptops, Apple Silicon M-series, and edge CPU nodes. Proprietary corporate intelligence never touches external foundation cloud APIs.

1.1 GB Footprint · Sub-Second Inference · Apache 2.0
CORE ARCHITECTURAL ADVANTAGES

Why Saransh? Five Core Innovations.

Engineered specifically to eliminate the verbosity, hallucinated references, and heavy resource footprints of general-purpose foundation models.

100% SHORT ADHERENCE

Length Control That Actually Holds

Ask for one or two sentences and get exactly one or two sentences. While the untuned base model spews an average of 252 words for short requests, Saransh produces concise, focused 24-word summaries.

Untuned Base: 252 wordsSaransh: 23.9 words
100% Target Match0% on base model
NUMERIC CONSTRAINTS

Explicit Numeric Word Budgets

Need ~120 words for an executive briefing card? Explicit numeric targets are trained into Saransh's weights across 35% of its entire training corpus, ensuring predictable UI container integration.

“Summarize the following text in about 120 words.”
35% of CorpusNumeric target alignment
VERIFIED

De-Hallucinated Citations

Purged 46%+ of training examples containing hallucinated citations (like “according to NYT”). Numbers, entities, and dates are strictly grounded in source passages.

0.04 Hallucination Score (-90% drop)
1.1 GB EDGE

No GPU Required

Runs locally on MacBooks (3.3s per summary on M3 Pro) and consumer edge nodes in 4-bit GGUF. Instant inference without renting cloud H100 clusters.

< 2.0 GB System RAM Needed
COMMERCIAL

Permissive Apache 2.0

Commercial enterprise usage without restrictive per-seat licensing, seat audits, or telemetry tracking. Free to deploy, bundle, and modify within your own products.

Open Weights & GGUF Quantizations
INTERACTIVE ARCHITECTURAL LAB

Empirical Length Control & Telemetry

Inspect how Saransh 1.7B enforces strict word budgets across 4 discrete synthesis regimes compared to untuned base models.

Active Prompt:“Summarize the following text in one or two sentences.”
Adherence: 100%
Input Document Source~245 words

The widespread adoption of foundation language models has transformed enterprise knowledge management, yet cloud-hosted models present persistent friction points: escalating inference costs, data residency vulnerabilities, and network round-trip latencies. To mitigate these obstacles, engineering teams are transitioning toward compact Small Language Models (SLMs) that execute directly on local CPU or edge hardware.

However, standard small models exhibit a notorious structural deficiency in summarization: an inability to adhere to length instructions. When prompted for a brief two-sentence executive brief, base models routinely produce rambling, multi-paragraph expositions averaging upwards of 250 words. In user-facing software, such unconstrained outputs shatter UI component boundaries, waste compute tokens, and dilute reader focus.

Saransh 1.7B addresses this limitation through targeted length-controlled alignment. Built on the Qwen3-1.7B architecture, the model underwent three successive data curation passes to eradicate phantom citations, normalize numeric entities, and couple explicit length buckets to distinct prompt phrasings. The resulting model delivers deterministic summaries matching the requested word count while maintaining strict factual alignment with the source text.

Source Context: 8,192 tokens maxFormat: ChatML
Saransh 1.7B Synthesized Output
24 words
Enterprise small language models allow organizations to perform deterministic, on-device summarization with strict word-budget adherence without exposing proprietary data to cloud APIs.
Qwen3-1.7B Base (Untuned) Response252 words
Saransh: 24w (100% adherence)Untuned base: 252w (+228w bloat)
Greedy Decoding (do_sample: False)Repetition Penalty: 1.05
M37Labs Enterprise AI Computing Lab
100% BEHIND FIREWALL · ZERO DATA EGRESS · PERMISSIVE APACHE 2.0

Summarization without the infrastructure overhead.

1.7B parameters. 1.1 GB quantized. Zero GPU required. Bring controllable, on-device document synthesis directly inside your sovereign enterprise workflows.