APXInf Open-Sourced: Robot Inference Engine Accelerates 10.7x on Thor Chip

icon MarsBit
Share
AI summary iconSummary
On-chain news from MarsBit reports that on September 15, 2026, Tsinghua University, Wuwen Xinqiong, and Shanghai Jiao Tong University open-sourced APXInf, an edge-side inference engine for embodied AI models. The engine runs 10.7x faster on Thor FP8 chips, reducing latency from 278ms to 26ms. APXInf is designed for real-time performance in robotics, enabling efficient model iteration and hardware adaptation. AI + crypto news highlights the growing synergy between AI and blockchain infrastructure.

The second week of September was bustling in the embodied AI community.

On the 8th, Songyan Power released the HERON-World Model; on the 9th, Zhiyuan unveiled both AGILE 2.0 and GE-Act 2.0 in one go; on the 10th, Unitree open-sourced UnifoLM-WLA-1.0, a 6B-parameter model trained on approximately 2,500 hours of real-world data, capable of managing 64 tasks with a single model; by the 15th, Songyan had additionally released HERON-CRA. In eight days, three robotics companies launched five models.

Meanwhile, robots are accelerating their entry into the physical world. On September 20, Qiyuan Robotics held a product launch event, officially releasing two personal robots priced from 19,999 yuan for the Q1 model. Prior to this, on September 10, UBTECH announced securing over 50 million yuan in overseas orders for products including the Walker C1 and U1. The capital market’s evaluation criteria are also shifting, now reassessing embodied AI companies based on churn-adjusted repeat purchase rates, operating net cash flow, and actual fulfillment costs—namely, the investments required for deployment, calibration, and maintenance after delivery.

Two threads converge, revealing a core pain point: How can we efficiently and stably deploy complex, powerful embodied large models onto robot hardware with extremely limited computing power? This is precisely the core value of edge-side inference engines.

On September 15, Tsinghua University, in collaboration with Wuwen Xinqiong and Shanghai Jiao Tong University, open-sourced APXInf, an edge-side inference engine designed for embodied models. It addresses two key questions simultaneously:

How can embodied models achieve usable inference speeds on robotic platforms in edge environments with limited compute, memory, and power?

As models continuously iterate and evolve, how can we build on-device optimization capabilities that adapt continuously and never fall behind?

The first question: APXInf delivered a set of numbers—without modifying the π0.5 model itself, end-to-end full-stack optimization reduced inference latency from 278ms to 26ms under the Thor chip’s FP8 configuration, achieving an end-to-end speedup of approximately 10.7x, with a frequency of 38.46Hz bringing robot control into the real-time range.

Bot

The answer to the second question is embedded in the build process of the APXInf repository, which transforms the previously expert-dependent model adaptation, optimization, and validation into a reusable workflow accessible to agents.

Bot

Project address: https://github.com/RLinf/APXinf-robo

Why does embodied AI require specialized edge-side inference optimization?

To answer this question, first understand where the reasoning engine is located within a robot.

An embodied agent typically consists of a main controller, a compute box, and peripherals. One control cycle proceeds as follows: the main controller collects observations, the compute box infers an action chunk, and sends it to the arm or base for execution, then moves to the next frame. The inference engine is embedded in the middle of this cycle—it determines how many milliseconds each inference takes, what control frequency in hertz can be achieved, and the size, heat output, and cost of the compute box.

Bot

This position imposes a set of fairly specific requirements on the engine: small batch sizes, real-time processing, low latency jitter, and stable invocation via WebSocket or ROS. These are precisely the areas where cloud-based inference frameworks are not strong.

The general inference framework lowers models layer by layer using a unified intermediate representation, then passes them to backends for executable code generation, with a single compilation stack covering as many models and hardware platforms as possible. In contrast, solutions like vLLM and SGLang are designed around cloud throughput. Both approaches have been highly successful, but their benefits rely on assumptions—diverse model types, large batch sizes, and ample scheduling flexibility—that are precisely the conditions not met in embodied edge scenarios.

Thus, on-device deployment of embodied models faces three practical bottlenecks.

Edge-side performance. On the edge, computational power, bandwidth, power consumption, and thermal dissipation are all severely constrained, yet a single inference must simultaneously handle multi-view perception, model forward propagation, and action generation while responding to the main controller at a steady rhythm. The memory bandwidth between edge modules and discrete GPUs differs by 4 to 8 times, and power consumption differs by 5 to 10 times. As a result, models that run smoothly in the cloud may perform significantly worse on Thor or Orin.

Human resources and time. Deploying a model to edge hardware is not as simple as copying and running it—it is a massive systems engineering endeavor. It requires a full end-to-end process, including low-level architecture adaptation, core operator compilation, precision quantization, and comprehensive software-hardware performance tuning and simulation validation. This process typically takes several weeks and relies on cross-disciplinary experts who understand both inference systems and operator optimization—professionals who are extremely scarce in any robotics company. More critically, even minor hardware changes can render all prior work obsolete: switching chips means completely redoing operator selection, memory layout, and pipeline scheduling.

Bot

Stability. Getting a demo to run and achieving long-term stable operation are two different things. The latter must handle continuous perception and control, limited resources, and coordination among multiple modules. What’s most concerning isn’t slowness, but rather sudden jitter, freezes, or loss of state control during operation.

When these three are stacked together, they form a product feature:

The total effort to deploy the model onto the ontology ≈ (one-time integration + one-time optimization) × number of ontology models × number of chip platforms × number of model iterations.

Each variable on the right is rapidly increasing: model variants are growing, chip diversity is expanding, and model iteration cycles are shrinking at an unprecedented pace. In fact, the five models introduced in just eight days at the start of this article directly reflect this reality. Simply expanding the team is not feasible—what’s needed is a reusable infrastructure layer that can push single inference to hardware limits and enable seamless integration of the next model or chip without starting from scratch.

APXInf is exactly what this layer needs.

What did APXInf do right to compress inference into the real-time window?

First, let’s look at the challenge: how to optimize inference efficiency on limited hardware.

APXInf does not prioritize a unified, general-purpose IR; instead, it focuses on building specialized execution paths tailored to model families. Model architecture, weight layout, memory allocation, operator fusion strategies, and execution order are directly encoded into the system. Only capabilities validated across multiple models are further abstracted into shared modules. This design enables deep optimization aligned with model structures and hardware characteristics, allowing targeted adjustments in operator selection, memory layout, and execution flow—reducing overhead introduced by accommodating generic, one-size-fits-all scenarios.

At runtime, retain only the control capabilities truly needed for on-device real-time inference:

Operator execution is completed using CUDA Graph for full graph capture and steady-state replay; kernel selection is generated and persisted by Autotune;

Memory is held by model layers as a fixed workspace with predetermined shapes, pre-allocated and stable addresses, reducing data movement along the hot path;

Scheduling is centered around real-time inference with small batches, without incorporating continuous batching or paged attention for high-throughput large batch processing.

By avoiding complex heuristics at runtime, we ensure execution paths are predictable, reproducible, and auditable.

This specialization extends all the way to the build phase: during compilation, the system queries the local GPU’s compute capability and compiles the kernel exclusively for that architecture; the source code for CUDA kernels, CUTLASS, and FlashAttention is bundled within the repository, eliminating the need for Docker or external framework dependencies and avoiding multi-gigabyte image sizes.

Ultimately, this full-stack optimization delivers a significant boost in efficiency. The optimization ladder on Orin illustrates this: starting from a baseline of 1300ms, successive improvements through torch.compile, Pipeline, Graph, Kernel, and Pruning reduce the time to just 119ms, achieving an overall speedup of over 10.9x. The same applies to Thor.

Bot

In terms of results, the official evaluation protocol consists of 50 episodes for each of the 10 LIBERO-10 tasks, with a fixed seed of 7 and a replan step size of 5, totaling 500 rollouts. Thor FP8 achieves a success rate of 92.2%, Thor BF16 achieves 92.8%, Orin BF16 achieves 92.0%, and the reference implementation π0.5 achieves 92.4%.

Bot

Beyond speed, stability matters just as much. APXInf’s底层 focuses on high-performance operators to squeeze out maximum performance; the core of the inference framework is developed in Rust, a systems language that provides low-cost, fine-grained control over system runtime, enabling superior scheduling and concurrency management. Meanwhile, Rust’s enforced RAII features and ownership system significantly narrow the risk surface of memory issues, supporting safer and more robust resource lifecycle management, with unsafe code strictly confined to fixed FFI boundaries. This system-level value in the middle layer ensures that hidden failures like dangling pointers and data races are eliminated before deployment; for a robot designed to operate continuously, it’s not enough to simply “run at 38Hz”—it must maintain 38Hz even after eight straight hours of operation.

Models are constantly evolving—how can inference optimization keep up?

Specialized paths deliver performance but also introduce new challenges: manually writing a path for each model family means that every time the model is updated or the hardware is upgraded, the path must be rewritten. If this work continues to rely on a small group of experts to complete manually, it will never keep pace with the iteration of embodied models.

Bot

APXInf's solution is to productize this engineering capability: transforming models for integration, preprocessing and postprocessing, ontology adaptation, performance optimization, and deployment validation—previously scattered across the expertise of a few specialists—into engineering workflows that code agents can understand, invoke, and continuously iterate upon.

In this entirely new workflow, the division of labor between humans and machines takes on a transformative pattern:

The agent is responsible for running the workflow: reading the PyTorch reference, generating an execution ledger, implementing the model's static path, performing operator-by-operator verification and regression, and finally completing autotuning, documentation, and iteration.

The person is responsible for setting standards: architectural boundaries and module responsibilities, kernel contracts and security guidelines, accuracy, performance, and acceptance criteria for tasks, as well as deciding when to extract commonalities into shared abstractions.

Bot

What is handed over to the Agent is no longer simple auxiliary tasks, but the most time-consuming implementation aspects that previously consumed experts’ time; humans evolve into the overarching "criteria" and "decision-makers."

Precisely because implementation can be delegated to Agents, the rigor of the verification system has become the new core moat. APXInf has established a three-layer safeguard:

End-to-end layered validation: From individual operators and layers to complete models, combined with cross-validation between Eager and Graph paths to ensure precise and controllable steps at every stage.

Fail-closed principle: Unsupported parameters or hardware result in immediate errors—never silently output incorrect results.

Model family isolation: Shared capabilities are only delegated to the Kernel layer, and any changes must undergo rigorous layered regression to contain risk.

What does this mean for users?

For embodied enterprises, it transforms inference optimization from a one-time project into a sustainable engineering capability. When modifying an in-house model, there’s no need to rejoin a waiting list for experts; switching to a different chip doesn’t require discarding previous optimization efforts; team size is no longer a hard constraint on integration speed. More importantly, expert knowledge becomes a reusable asset for the entire team, rather than relying on individual intuition.

For the industry, it offers a new approach to infrastructure development. In the past, evaluating an inference framework meant looking at how many models it supported, how many hardware platforms it adapted to, and how comprehensive its operator library was—underlying assumptions being that onboarding costs were fixed, so broader coverage was always better. But as onboarding costs themselves begin to decline, what truly matters becomes: how long it takes to onboard a new model or deploy a new ontology. This measures how quickly a robotics company can turn new models into actual production capacity.

From Mizar to APXInf, the edge inference tech stack extends into the physical world.

APXInf is not an experimental project started from scratch, but rather the inevitable outcome of Wuwen Xinqiong's long-term accumulation in edge-side inference.

APXInf's technical foundation is directly inherited from Mizar, an inference acceleration engine designed for intelligent terminals. Even during the Mizar phase, Wuxin Qiong had already developed core capabilities around AI PCs, AI boxes, and all-in-one devices, including heterogeneous hardware adaptation, local model deployment, inference acceleration, memory optimization, and low-power operation. This technology stack has not only enabled local model sizes to increase several-fold and inference performance to double, but also delivered an 18% improvement in intelligence under the same computational resources.

More importantly, this technology path has moved beyond the experimental stage and has undergone rigorous mass production validation. In February 2025, Wuwen Xinqiong entered into a deep collaboration with Lenovo, integrating the Mizar engine into Lenovo’s next-generation AI PCs, and by November of the same year, achieved a pre-installation partnership exceeding ten million AI PCs. APXInf is the next evolution built upon this foundation.

The workloads of AI PCs and embodied edge devices are different, but both face the same constraints: achieving usable inference speeds within limited computational power, memory, and power consumption.

To address the new challenges posed by embodied scenarios, APXInf has extended its technology: whether tackling more complex multimodal execution pipelines of VLA models or meeting millisecond-level real-time control demands for robots, APXInf fully unleashes hardware performance through deep underlying optimizations. Meanwhile, to keep pace with rapid model iteration, it introduces an engineering workflow designed specifically for Agent development. This not only enables a seamless migration of existing edge capabilities but also achieves deep adaptation tailored for the embodied industry.

Another consideration in the ecosystem layout is the seamless integration of an "engineering闭环." APXInf chose to launch within the RLinf open-source ecosystem because model training and ontology reasoning are inherently two stages on the same engineering pipeline.

RLinf covers reinforcement learning training, evaluation, and real robot workflows; however, for trained embodied models to be truly deployed on hardware, a dedicated inference engine designed for small batches, low latency, and limited power consumption is required. APXInf picks up where RLinf leaves off, extending its capabilities from training and evaluation to edge-side inference and deployment, enabling developers to complete the entire workflow—from training and validation to on-device execution—within a unified ecosystem.

Next, open-source collaboration

Today, APXInf has successfully completed integration with π0-fast, GR00T, and Qwen Drive, further expanding model support to diverse embodied scenarios such as robotic manipulation and mobile agents.

On the hardware side, AMD and domestic chip backends have also been incorporated into the roadmap. In the future, we will support more embodied models running efficiently and stably on various edge devices.

For embodied intelligence to truly enter the physical world, the model, the ontology, and this intermediate layer of infrastructure are all indispensable. Faced with the massive engineering effort required for integration and verification, this is precisely why APXInf chose to open-source and collaborate共建.

If you're deploying embodied models like π0.5 onto robots, and you have access to an RTX 4090, Jetson Thor, or Orin—or if you'd like to integrate new models or chips into APXInf, or contribute to building the model integration and agentic deployment workflows—feel free to get started, submit an issue, or contribute a PR.

GitHub: https://github.com/RLinf/APXinf-robo

Quick Start: https://github.com/RLinf/APXinf-robo#build-apxinf-robo

Thank you

The development of APXInf was inspired by the following outstanding open-source projects. Wuwen Xinqiong extends sincere gratitude to the communities behind these projects for their open-source spirit and technical contributions.

FasterTransformer:

https://github.com/NVIDIA/FasterTransformer

Tensor-LLM:

https://github.com/NVIDIA/TensorRT-LLM

llama.cpp:

https://github.com/ggml-org/llama.cpp

vLLM:

https://github.com/vllm-project/vllm

sgLang:

https://github.com/sgl-project/sglang

FlashRT:

https://github.com/flashrt-project/FlashRT

This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by Machine Heart.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.