
Summarization that follows your length.
Generalist LLMs ignore word count constraints and leak sensitive corporate data to external clouds. Saransh 1.7B is an edge-optimized sovereign model built for strict length budgets and calibrated synthesis on local hardware.
Small model. Strict word budgets. Zero cloud egress.
Saransh (सारांश, Sanskrit for “summary”) was engineered to solve the primary failure mode of small models: uncontrolled verbosity. Untuned base models spew 250+ rambling words when asked for brief summaries, hallucinate phantom sources, and exhaust inference budgets.
One to Three Pages Uncut
Paste dense financial reports, earnings filings, or technical transcripts. With an 8,192 trained token window, documents are ingested as cohesive narratives without breaking paragraphs into fragmented RAG chunks.
Strict Length Adherence
Request a two-sentence headline and get exactly two sentences. Choose from four discrete length regimes or provide exact word count budgets to fit strict UI layout constraints without truncation artifacts.
Zero Cloud Egress
Quantized down to 1.1 GB in 4-bit GGUF. Runs natively on standard laptops, Apple Silicon M-series, and edge CPU nodes. Proprietary corporate intelligence never touches external foundation cloud APIs.
Why Saransh? Five Core Innovations.
Engineered specifically to eliminate the verbosity, hallucinated references, and heavy resource footprints of general-purpose foundation models.
Length Control That Actually Holds
Ask for one or two sentences and get exactly one or two sentences. While the untuned base model spews an average of 252 words for short requests, Saransh produces concise, focused 24-word summaries.
Explicit Numeric Word Budgets
Need ~120 words for an executive briefing card? Explicit numeric targets are trained into Saransh's weights across 35% of its entire training corpus, ensuring predictable UI container integration.
De-Hallucinated Citations
Purged 46%+ of training examples containing hallucinated citations (like “according to NYT”). Numbers, entities, and dates are strictly grounded in source passages.
No GPU Required
Runs locally on MacBooks (3.3s per summary on M3 Pro) and consumer edge nodes in 4-bit GGUF. Instant inference without renting cloud H100 clusters.
Permissive Apache 2.0
Commercial enterprise usage without restrictive per-seat licensing, seat audits, or telemetry tracking. Free to deploy, bundle, and modify within your own products.
Empirical Length Control & Telemetry
Inspect how Saransh 1.7B enforces strict word budgets across 4 discrete synthesis regimes compared to untuned base models.
The widespread adoption of foundation language models has transformed enterprise knowledge management, yet cloud-hosted models present persistent friction points: escalating inference costs, data residency vulnerabilities, and network round-trip latencies. To mitigate these obstacles, engineering teams are transitioning toward compact Small Language Models (SLMs) that execute directly on local CPU or edge hardware.
However, standard small models exhibit a notorious structural deficiency in summarization: an inability to adhere to length instructions. When prompted for a brief two-sentence executive brief, base models routinely produce rambling, multi-paragraph expositions averaging upwards of 250 words. In user-facing software, such unconstrained outputs shatter UI component boundaries, waste compute tokens, and dilute reader focus.
Saransh 1.7B addresses this limitation through targeted length-controlled alignment. Built on the Qwen3-1.7B architecture, the model underwent three successive data curation passes to eradicate phantom citations, normalize numeric entities, and couple explicit length buckets to distinct prompt phrasings. The resulting model delivers deterministic summaries matching the requested word count while maintaining strict factual alignment with the source text.

Summarization without the infrastructure overhead.
1.7B parameters. 1.1 GB quantized. Zero GPU required. Bring controllable, on-device document synthesis directly inside your sovereign enterprise workflows.

