DATAFLOW-DRIVEN GPNPU

Redesigning compute, interconnect and software around LLM inference dataflow.

The edge-inference bottleneck is not just peak compute. Weight capacity, effective bandwidth, non-matrix operators and model evolution require compute, data movement and compiler scheduling to be co-designed.

01

FusionCore

A multi-precision compute IP that organizes matrix, vector and special-function paths, on-chip memory and data movement into composable modules for edge LLMs, low-bit formats and precision fallback.

02

ScaleNoC

A dataflow interconnect for multiple compute cores, on-chip memory and external interfaces, configurable by topology, links and bandwidth, with optimization for sparse compute and MoE weight movement.

03

TileFlow

Compiler, scheduling and runtime arrange tiling, layout, movement and synchronization around operator shapes and on-chip resources, improving effective bandwidth through fusion and compute/movement overlap.

ARCHITECTURE

Three layers designed together, not isolated compute.

Compute formats, interconnect and software scheduling are jointly optimized around the model, forming three core assets that can be licensed separately or combined.

Compute

Multi-precision matrix/vector compute, low-bit weight formats and on-chip buffering co-designed.

Interconnect

Configurable dataflow organizes movement among compute, memory and external interfaces.

Software

Connects models, operators and hardware execution to reduce deployment and migration cost.

MODEL SUPPORT

Built for dense models, sparse MoE and multi-agent workloads.

One architecture supports different activation patterns and model structures for real edge workloads.

01

Dense models

Controls capacity and memory cost with multi-precision compute and programmable dataflow.

02

Sparse MoE

Organizes compute and weight access by active experts to reduce active weight traffic per token.

03

Multi-agent

Supports multiple simultaneous inference tasks on one device.

VALIDATION

Validated progressively from RTL to prototype.

Current engineering status is subject to company confirmation; performance, power and interface details will follow silicon and system testing.

RTL + DV

Core compute modules have completed engineering-level development and remain under validation.

FPGA validation

Qwen-8B completed end-to-end validation on the FPGA prototype.

Chip integration

The first-generation product is advancing through system integration, back-end and tape-out preparation.

Silicon evaluation

Evaluate product capability through silicon testing and customer workloads.

SYSTEM VIEW

Keep business on the host and focus the on-chip system on inference.

The external host keeps running the OS and applications; the on-chip management CPU schedules tasks while compute tiles handle matrix, vector and data movement.

GPNPU system tile architecture
The GPNPU system combines compute tiles, on-chip interconnect, memory interfaces and management control; configuration is project-specific.

NEXT

Move IP capability into deployable products.

VLX-64 is offered in M.2 and USB engineering-module formats; VLX-512 will target larger capacities, workstations and edge systems.

View products