FPGA compute for AI inference

Reconfigurable inference, datacenter to edge

Dataflow accelerators with deterministic latency and custom precision — for datacenters that aren't power-bounded, and for drones and cars at the edge.

Why FPGA

The chip adapts to the model

Reconfigurable dataflow

The chip becomes the model. Build a custom dataflow pipeline per network instead of forcing it through a fixed instruction set — and reprogram the fabric when the model changes.

Deterministic latency

No batching games, no scheduler jitter. Inference runs in hardened logic with bounded, repeatable per-inference latency — the number real-time and safety-critical loops actually need.

Custom precision

INT8, INT4, binary, or bespoke block-float. Quantize to exactly the precision your accuracy budget allows and spend the freed silicon on throughput.

Longevity, no lock-in

10–15 year silicon lifecycles and an open toolchain. Retarget the same design from a datacenter card to an edge module without rewriting your stack.

Datacenter · not power-bounded

Optimize for latency and throughput, not watts

When you own the rack and the power budget, the constraint isn't efficiency — it's tail latency and throughput-per-slot. A reconfigurable dataflow fabric serves inference with bounded latency and precision tuned to the model, and reprograms as your models change.

SL-D1

Meridian D1

Datacenter FPGA inference accelerator

  • AMD Versal-class adaptive SoC
  • PCIe Gen5 ×16 · 32 GB HBM2e
  • INT8 / INT4 / BF16 + custom precision
  • Deterministic low-latency dataflow · ~225 W

SL-D2

Meridian D2-X

Dual-FPGA accelerator for maximum throughput

  • Two D1-class fabrics on one full-height card
  • Highest throughput-per-slot
  • For LLM, recsys, and video inference serving
  • ~350 W · passive datacenter airflow
Contact sales →

Edge · drones & automobiles

Real-time inference where the data is born

At the edge the pressures invert: watts, size, and a hard real-time deadline. FPGA modules deliver deterministic, functional-safety-grade latency in a handful of watts — running fully offline for drones, robots, and vehicles when the link is denied.

SL-E1

Talon E1

Edge FPGA inference module — drones & robotics

  • Kria-class adaptive SOM · ~15 W
  • 77 × 60 mm · sub-ms deterministic latency
  • Multi-camera sensor fusion
  • Runs fully offline — no cloud round-trip

SL-E2

Sentinel E2

Automotive FPGA inference module — ADAS / AD

  • Functional-safety target (AEC-Q100 / ISO 26262)
  • Extended temperature range
  • Camera + radar + lidar fusion
  • Deterministic real-time perception
Contact sales →

Model-Ops

Keep every deployment current

Quantize → compile with Vitis AI → deliver over the air → hot-swap and monitor. The same signed, versioned, auto-reverting pipeline serves a rack of datacenter cards or a fleet of edge modules in the field.

Read the full architecture →

Applications · datacenter

Serving inference at scale

LLM & generative serving

Low-latency token generation and embeddings at custom precision — without the power ceiling of a GPU farm.

Recommendation & ranking

High-QPS, tight-tail-latency inference for feeds, ads, and search.

Video & vision analytics

Real-time decode-and-infer across thousands of concurrent streams.

Applications · edge

Autonomy in the physical world

Drones & robotics

On-board perception, navigation, and targeting that keeps working when the link is denied.

Automotive · ADAS & AD

Deterministic, functional-safety-grade sensor fusion for driver assistance and autonomy.

Autonomous machines

Agriculture, mining, and industrial platforms running inference where connectivity is poor or absent.