Skip to content

FS-Attention

Hybrid attention inference IP

Native support for hybrid-attention large language models.

FS-Attention Hybrid attention inference IP
Fig. 1 · FS-Attention Hybrid attention inference IP

Data flow

Native support for hybrid-attention large language models.

token全局注意力头softmax滑窗注意力头localKV 缓存on-chip下一 tokenQwen3 / 3.5 · SmolLM2
Fig. 2 · FS-Attention data flow

Native support

Qwen3 / 3.5 (0.6B to 9B) and SmolLM2

Deployment

ASIC tape-out and FPGA prototyping, with interface and system adaptation per project.

Licence includes

Inference core IP, FPGA evaluation board, model mapping and interface adaptation, design-in support.

IP coverage, interfaces, resource usage, performance targets and licensing terms are evaluated per target device and model.

Ask about FS-Attention

Tell us the target device, model and interfaces. We reply with IP coverage and an evaluation plan.

Ask about IP licensing