// SILICON IP

License the IP.
Build your own silicon.

ExSLerate is the AI accelerator IP family that powers the Krsna SoC — and is licensable as standalone IP to silicon vendors, OEMs, and chip startups building their own AI hardware. Four configurations today (Lite to Apex). Two patented engines inside V2: Dynamic Neural Compression and the Infinite Series Engine. Two generations ahead on the roadmap. ARM-style licensing model for AI silicon.

Current generation
V2
Variants shipping
4
Generations on roadmap
4
Foundational year
2019

Buy the chip. Or license the IP.

Most AI silicon companies sell you one thing. We sell you two — because the buyers are different and the economics are different. Krsna SoC is the finished chip. ExSLerate is the IP behind it. Pick the model that matches your business.

// FINISHED CHIP

Krsna SoC →

For OEMs and device-makers buying a finished, production-ready AI chip. Per-chip pricing. Reference designs included. Lite for wearables, Apex for data center.

Buyer: OEMs, device-makers · Sale: per-chip · Margin: hardware
// LICENSABLE IP

ExSLerate IP

For silicon vendors and chip startups designing their own SoCs. Licensable IP — RTL, compiler stack, integration support. ARM-style commercial model: license fee + per-unit royalty.

Buyer: Silicon vendors · Sale: license + royalty · Margin: IP

Four generations. Edge to data center.

V1 was the 2019 microprocessor challenge winner. V2 is FPGA-validated today across the Krsna SoC configurations. V3 climbs to the deskside appliance tier, where the comparison is a workstation GPU like the RTX 6000 Ada. V4 enters the data center as an inference blade benchmarked against the H100 — serving, not training.

2019
V1 · Embedded
Foundational IP. Won India's Microprocessor Challenge — Ranked #1 of 30 finalists. Established the architectural patterns that V2 builds on.
2024–2026
V2 · Endpoint + robotics
4-configuration family — Lite (M64, always-on wearables) → Apex (M4096, robotics + automotive). Native FP8 (E4M3) and FP4 precision. FPGA-validated on AMD Xilinx ZCU106 and Kria KR260. Three innovations inside: the patented Tensulator spherical buffer, the SL Tensor Codec (lossless FP8→FP8 payload reduction), and the Special Function Unit for non-linear math on-chip. 40–50% less memory traffic and RAM at peak context — enough to run Llama 3.1 8B on an 8 GB endpoint.
Next gen
V3 · SOHO deskside appliance · 27B-class inference
Tensor Codec Gen-2 — 60% memory usage and traffic reduction. ~1,000 TOPS FP8 / BF16 dense at ~75 W, with ~800 GB/s effective bandwidth over commodity 24 GB GDDR6. Runs a 27B model at 128k context on a deskside appliance — 4× the usable context of an RTX 6000 Ada at a quarter of the power, and 2.5× its efficiency. Built for regulated businesses that need private 27B-class reasoning on-premise.
Future
V4 · Data-center inference blade · 70B-class serving
Tensor Codec Gen-3 — 70% memory usage and traffic reduction. ~3,000 TOPS FP8 / BF16 dense (6,000 FP4) at ~250 W, with ~2.2 TB/s effective bandwidth over commodity GDDR6 instead of supply-constrained HBM. Roughly 1.5× H100 compute at 4× the efficiency, in a standard rack. Built for serving 70B-class models where the economics are tokens per dollar and tokens per watt — not peak TOPS.
Each generation targets a higher compute tier. V3 and V4 are co-designed with EdgeFlow — so when the runtime absorbs new model architectures, the silicon is already ready for them.
V1
2019
SHIPPED

Embedded

Foundational IP. Won India's Microprocessor Challenge — Ranked #1 of 30 finalists. Established the architectural patterns that V2 builds on.

MeitY MPC ranking
#1
Target
Embedded edge
V2
2024–2026
CURRENT

Endpoint + robotics

4-configuration family — Lite (M64, always-on wearables) → Apex (M4096, robotics + automotive). Native FP8 (E4M3) and FP4 precision. FPGA-validated on AMD Xilinx ZCU106 and Kria KR260. Three innovations inside: the patented Tensulator spherical buffer, the SL Tensor Codec (lossless FP8→FP8 payload reduction), and the Special Function Unit for non-linear math on-chip. 40–50% less memory traffic and RAM at peak context — enough to run Llama 3.1 8B on an 8 GB endpoint.

Configurations
4
MAC range
M64 → M4096
Peak compute · Apex
22 TOPS FP8
Efficiency
~4.4 TOPS/W
Apex TDP
< 5 W
Memory + traffic cut
40–50%
V3
Next gen
ROADMAP

SOHO deskside appliance · 27B-class inference

Tensor Codec Gen-2 — 60% memory usage and traffic reduction. ~1,000 TOPS FP8 / BF16 dense at ~75 W, with ~800 GB/s effective bandwidth over commodity 24 GB GDDR6. Runs a 27B model at 128k context on a deskside appliance — 4× the usable context of an RTX 6000 Ada at a quarter of the power, and 2.5× its efficiency. Built for regulated businesses that need private 27B-class reasoning on-premise.

Tensor Codec
Gen-2
Memory + traffic cut
60%
Peak compute · FP8
~1,000 TOPS
Efficiency
~13 TOPS/W
Power envelope
~75 W
Context · 27B
128k tokens
V4
Future
ROADMAP

Data-center inference blade · 70B-class serving

Tensor Codec Gen-3 — 70% memory usage and traffic reduction. ~3,000 TOPS FP8 / BF16 dense (6,000 FP4) at ~250 W, with ~2.2 TB/s effective bandwidth over commodity GDDR6 instead of supply-constrained HBM. Roughly 1.5× H100 compute at 4× the efficiency, in a standard rack. Built for serving 70B-class models where the economics are tokens per dollar and tokens per watt — not peak TOPS.

Tensor Codec
Gen-3
Memory + traffic cut
70%
Peak compute · FP8
~3,000 TOPS
Efficiency
~12 TOPS/W
Power envelope
~250 W
Memory
Commodity GDDR6
// V2 · CURRENT GENERATION

Four configurations. One IP family.

ExSLerate V2 ships across four Krsna SoC configurations — Lite (M64) for always-on wearables, Pulse (M256) for smartwatches and smart speakers, Surge (M1024) for drones and aerial platforms, and Apex (M4096) for robotics and automotive. Same IP family, scaled across the full edge-to-robotics envelope.

VariantIPMAC countPrecisionTarget
Krsna ApexM40964,096INT4 · FP8Robotics · Automotive · Industrial
Krsna SurgeM10241,024INT4 · FP8Drones · Aerial · Light edge
Krsna PulseM256256INT4 · FP8Smartwatch · Smart speaker
Krsna LiteM6464INT4Always-on wearables · Hearables

Climbing the NVIDIA stack.

V2 is shipping today across endpoint and robotics. V3 climbs to the SOHO deskside tier — Tensor Codec Gen-2 fits a 27B model at 128k context onto a single 24 GB GDDR6 card at ~75 W. V4 enters the data center with Tensor Codec Gen-3 at ~3,000 TOPS FP8 and ~250 W. Each generation pushes the reduction ratio further; each unlocks a higher tier of model size on commodity memory instead of HBM.

V3
Next gen · Roadmap

SOHO deskside appliance · 27B-class inference

Tensor Codec Gen-2 cuts memory use and traffic by 60%. ~1,000 TOPS FP8 at ~75 W, with ~800 GB/s effective bandwidth over a single 24 GB GDDR6 card — half the RAM and a third of the bus width of the standard requirement, at 128k context on a 27B model.

Tensor Codec
Gen-2
Memory + traffic cut
60%
Peak compute · FP8
~1,000 TOPS
Power envelope
~75 W
V4
Future · Roadmap

Data-center inference blade · 70B-class serving

Tensor Codec Gen-3 targets a 70% cut in memory use and traffic. ~3,000 TOPS FP8 (6,000 FP4) at ~250 W with ~2.2 TB/s effective bandwidth — roughly 1.5× H100 compute at 4× the efficiency, on commodity GDDR6 instead of supply-constrained HBM.

Tensor Codec
Gen-3
Memory + traffic cut
70%
Peak compute · FP8
~3,000 TOPS
Power envelope
~250 W

Detailed throughput, latency, power, and per-configuration benchmarks are released under NDA on an engagement basis.

Not just RTL. The full stack.

A typical AI IP license drops you RTL and tells you to figure out the rest. ExSLerate licensees get the IP plus the runtime that's already optimized to run on it — because we built both layers together.

// LAYER 01

RTL deliverables

Synthesizable Verilog RTL for the selected variant. Verification suite. Integration documentation.

// LAYER 02

Compiler stack

CORE compiler pre-tuned for the licensed variant. Quantization, kernel scheduling, op fusion — all included.

// LAYER 03

Runtime

EdgeFlow inference engine that runs out-of-the-box on your silicon. 193 model architectures pre-supported. Built on IREE / MLIR — open frontends, no vendor lock.

// LAYER 04

Integration support

Engineering team available for SoC integration, customization, and tape-out support. Not a hands-off license.

// PROOF

ExSLerate has been getting jury-validated since 2019.

2019 — ExSLerate V1 ranked #1 of 30 finalists in MeitY's India Microprocessor Challenge. Foundational silicon recognition that seeded the IP family.

2023 — Aegis Graham Bell Award for the chip program. Selected into MeitY C2S — 1 of 13 companies in India's flagship semiconductor program.

2024 — Selected into Qualcomm QSMP as 1 of 2 cohort companies — industry-partner validation from the chip leader.

2025 — Co-development partnership with Brandworks Technologies announced. First wave of co-developed AI hardware planned for 2026.

// LET'S BUILD

License ExSLerate. Build your silicon.