VELOXIS AI Official Website

DATAFLOW-DRIVEN GPNPU

Built for edge LLM inference,
rethinking the dataflow.

Hangzhou VELOXIS AI Technology Co., Ltd. builds a full-stack GPNPU architecture with FusionCore compute, ScaleNoC interconnect and TileFlow compiler/runtime, offering licensable IP and VLX accelerator cards for existing devices.

64 TFLOPs FP4VLX-64 design peak
MXFP4 + MoENative support for edge LLM workloads
VELOXIS AI GPNPU architectureFusionCore programmable compute core, ScaleNoC interconnect, TileFlow compiler and runtime.FusionCorePROGRAMMABLE COMPUTETILEFLOW / RUNTIMELLAMA.CPPVLLM / QUANTSCALENOC / DATAFLOWDENSE + SPARSEMoE
COMPUTE / INTERCONNECT / SOFTWAREVELOXIS AI
01 / COMPUTEFusionCoreMulti-precision matrix, vector and on-chip memory co-design
02 / INTERCONNECTScaleNoCOrganizes compute cores and memory around dataflow
03 / SOFTWARETileFlowCompiler, scheduling and runtime
04 / VALIDATIONQwen-8BCompleted end-to-end FPGA prototype validation

WHY NOW

The edge-LLM bottleneck is
not peak compute.

Token generation is constrained by capacity and bandwidth, while model architectures keep evolving. Edge systems need compute, memory and software scheduling organized around real dataflow.

01

Compute and memory co-design

Starting from real inference dataflow, compute, load and write-back overlap to improve effective bandwidth utilization.

02

Multi-precision and MoE

Supports edge-oriented low-bit formats, with optimization for MoE expert scheduling and weight movement.

03

No platform replacement

Adds local inference to existing industrial PCs, compute boxes and terminals without redesigning the platform.

BUSINESS LINES

From licensable IP,
to standard accelerator cards.

Both business lines share one compute, interconnect and software stack: one serves chip and subsystem customers, the other provides standard products to device makers.

PRODUCT

GPNPU Chips & Accelerator Cards

From IP and subsystems to the VLX product family. VLX-64 targets edge inference in M.2 and USB engineering-module formats.

  • VLX-64:M.2 8 / 16 / 32GB
  • USB engineering module: 16GB
  • VLX-512: future high-capacity product roadmap
IP

AI Compute & Interconnect IP

Provides FusionCore compute IP, ScaleNoC interconnect IP, subsystems and custom design services to chip companies.

  • Available as standalone licenses or combined delivery
  • Configured by model, precision, interfaces and system scale
  • TileFlow provides compiler and runtime support

USE CASES

Bring large-model capability
to existing devices.

Extend AI across industrial PCs, compute boxes, smart terminals and robots while keeping the existing CPU, OS and application stack.

01

Industrial PCs

Run knowledge Q&A, document retrieval and voice interaction together.

02

Compute boxes

Add deployable local inference to industry devices.

03

Smart terminals

Delivers real-time interaction under power and cost constraints.

04

Robots

Run vision, language, planning and control models in parallel.

LATEST

Company updates

View all news

Qwen-8B completes end-to-end validation on the FPGA prototype

Technology update

VELOXIS AI closes a tens-of-millions RMB angel round

Company

First-generation GPNPU advances toward integration and tape-out

Technology

WORK WITH US

Build real edge-AI products with us.

We welcome device makers, chip and software teams, and people who want to build the GPNPU with us.

Business contact: info@veloxis-tech.com

Partner with us