Cheaper AI compute for
inference and training
FPGA systems that cut the cost of running — and training — AI models. Reconfigurable dataflow at custom precision, from datacenters that aren't power-bounded to drones and cars at the edge.
Scorpion Labs · Dendi Suhubdy · San Francisco · 2026
The Problem
AI runs on the most expensive compute ever sold
Almost all AI inference and training runs on GPUs — priced at monopoly margins, supply-constrained, and power-hungry. Every token served and every training run is billed against 70%+ hardware gross margins and a fixed dataflow that leaves silicon idle.
70%+
gross margin on the GPUs that run AI — you pay a monopoly tax on every FLOP, for inference and training alike
$/token ↑
inference cost scales with every user and never stops; the bill compounds as adoption grows
Power-bound
the largest datacenters now hit a megawatt ceiling, not a demand ceiling — watts are the constraint
Idle silicon
a fixed GPU pipeline runs your model through hardware it doesn't need, burning power on overhead
Why FPGA
The chip adapts to the model
A GPU makes the model fit the hardware. An FPGA makes the hardware fit the model — which is where the cost comes out.
Reconfigurable dataflow
fabric advantage
The chip becomes the model. Build a custom pipeline per network instead of forcing it through a fixed instruction set — and reprogram the fabric when the model changes.
Custom precision
fabric advantage
INT8, INT4, FP8, binary, or bespoke block-float. Quantize to exactly the precision the accuracy budget allows and spend the freed silicon on throughput.
Deterministic latency
fabric advantage
No batching games, no scheduler jitter. Bounded, repeatable per-inference latency in hardened logic — the number real-time and safety-critical loops need.
Longevity, no lock-in
fabric advantage
10–15 year silicon lifecycles and an open toolchain. Retarget the same design from a datacenter card to an edge module without rewriting the stack.
The Economics
The same result, for a fraction of the cost
3–5×
lower $/token on inference vs a comparable GPU
2–3×
lower total cost per training run (TCO, suitable workloads)
3–4×
better performance-per-watt
Where the savings come from
- Custom precision. Run at INT4 / FP8 / bespoke formats — fewer transistors switched per operation than a general FP16/FP8 GPU path.
- Dataflow, not fetch-execute. Weights stream through a hardened pipeline; no instruction scheduler, no idle SMs, no cache thrash burning power on overhead.
- Right-sized fabric. Buy exactly the compute the workload needs — no monopoly margin baked into a general-purpose part.
- Runs cooler. Lower watts per result means lower power and cooling cost — the other half of datacenter TCO.
The Product
Two families, one fabric
The same reconfigurable dataflow and toolchain, packaged for two very different power envelopes.
Datacenter · not power-bounded
Meridian D1 · D2-X
- PCIe Gen5 FPGA accelerator cards, up to 32 GB HBM2e
- Deterministic low-latency inference serving at scale
- Optimize for tail latency and throughput-per-slot, not watts
Edge · drones & automobiles
Talon E1 · Sentinel E2
- ~15 W adaptive modules with sub-ms deterministic latency
- Multi-camera / radar / lidar sensor fusion, fully offline
- Functional-safety target for automotive (AEC-Q100 / ISO 26262)
Inference & Training
Land on inference. Expand into training.
Inference — the beachhead
today
Inference is the larger and faster-growing share of AI spend, it never stops running, and it is where custom precision pays off most. That is where cheaper $/token turns directly into margin — for both datacenter serving and edge autonomy.
Training — the expansion
next
The same fabric scales into FPGA training clusters — cheaper pre-training and fine-tuning for workloads that benefit from custom precision and reconfigurable dataflow. It widens the market from “run the model” to “build the model.”
Deployed at the edge
Running where the data is born
Industrial / Robotics
Vision QC · predictive maintenance · on-machine perception
Aerial / Drone
Onboard nav & perception, fully offline in the air
Automotive / ADAS
Deterministic sensor fusion for driver assistance & autonomy
Why Now
The window is open
Demand is exploding, supply is constrained, and the software finally favors reconfigurable hardware. Low-precision numerics are mainstream and model architectures keep moving — so the flexibility of an FPGA stops being a tax and becomes an edge.
$106B → $255B
AI inference market 2025→2030 (~19% CAGR) — now larger and faster-growing than training
GPU-scarce
compute is supply-constrained and priced at monopoly margins — buyers are actively hunting alternatives
INT4 / FP8
low-precision inference and training are now standard — exactly the regime FPGAs win in
Model churn
architectures change faster than an ASIC tapes out — reconfigurable silicon keeps up where fixed silicon can't
Positioning
Reconfigurable and cheap
GPUs are flexible but expensive; ASICs are cheap but frozen. FPGAs are the middle that moves — and that's where cost-efficient AI compute lives.
| Approach | Strength | The catch — our opening |
|---|---|---|
| GPU — NVIDIA / AMD | General-purpose, mature software, runs anything. | Monopoly margins, power-hungry, fixed dataflow — you pay for silicon you don't use. |
| Inference ASIC — Groq / Cerebras / d-Matrix | Excellent perf/W and $/token — for the model they were designed around. | Frozen at tape-out. When the architecture moves, the silicon can't follow. |
| FPGA — ScorpionLabs | Reconfigurable dataflow at custom precision — cheap and adaptable, inference and training. | Ours to make easy: standard frameworks in, compiled fabric out — no new religion to adopt. |
Applications
From the rack to the road
Datacenter
- LLM & generative serving. Low-latency tokens and embeddings at custom precision — without the power ceiling of a GPU farm.
- Recommendation & ranking. High-QPS, tight-tail-latency inference for feeds, ads, and search.
- Training clusters. Cheaper pre-training and fine-tuning for workloads suited to reconfigurable dataflow.
Edge
- Drones & robotics. Onboard perception and navigation that keeps working when the link is denied.
- Automotive · ADAS & AD. Deterministic, functional-safety-grade sensor fusion for assistance and autonomy.
- Autonomous machines. Agriculture, mining, and industrial platforms running inference where connectivity is poor.
The Takeaway
Make AI compute cheap.
Inference and training.
Reconfigurable FPGA systems — from datacenters that aren't power-bounded to drones and cars at the edge. The cost of an ASIC, the flexibility of a GPU.
Scorpion Labs · Dendi Suhubdy · San Francisco
Scorpion