Technical notes on AI systems

Building software where agents are inside the system, not just beside it.

My current work sits between GPU infrastructure, cloud runtimes, and agent engineering. The question I keep returning to is simple: what changes when an application is designed with agents as durable parts of the architecture?

Work and interests

Less resume, more technical workbench.

AI Infrastructure Architect with 15+ years across high-performance computing, distributed systems, advanced networking, and large-scale AI infrastructure.

I use this site as a technical notebook for systems that move AI from model execution into persistent software: GPU clusters, remote PCIe/CXL fabrics, Kubernetes agent fleets, memory systems, and human-guided agent workflows.

The professional scope is intentionally layered. The lower layer asks how compute and memory move. The middle layer asks how agents run with state, identity, and audit. The upper layer asks how humans and agents collaborate on real engineering work.

01 Systems substrate GPU, RDMA, NVLink, PCIe, CXL, MoE inference, memory pressure.
02 Agent runtime OKE/FSS, durable workspaces, MCP control planes, private execution.
03 Engineering loop Deterministic issue graphs, memory, context, harnesses, review.

Three layers

A practical map for agent-native software.

Use the tabs to move between the layers. This is the main narrative of the site: infrastructure first, governed agent runtime second, experimentation third.

GPU systems and clusters

Cloud GPU systems, RDMA networking, NVLink environments, disaggregated PCIe/CXL infrastructure, memory pooling, and accelerator composition. This layer is about the physical and logical substrate where large AI workloads run.

RDMA NVLink PCIe/CXL MoE inference
systems notelayer 01
> substrate: compute + memory + fabric > pressure: GPU memory is scarce; movement must be scheduled > output: architecture choices that survive real workloads

Recent writing

Architecture notes, not marketing posts.

The writing is grouped by the same layers: GPU systems, agent-native infrastructure, agent engineering, and cloud architecture.

Selected projects

Projects as evidence for the writing.

The homepage keeps projects concise. Each one points to a system idea: control planes, memory pressure, remote fabrics, or heterogeneous execution.

Show more projects
  • vLLM MoE CPU Offload: GPU-native MoE offload using CPU host memory as the expert-weight bank. Repo
  • llama.cpp-MoE: Router-aware GPU expert slots for local MoE inference under constrained GPU memory. Repo
  • GPU-native Scheduler for GPU Computing: Patent-submitted scheduling approach for GPU resource management.
  • SuperKernel for SuperPod GPU Clusters: Patent-submitted Jupyter kernel architecture for large GPU fabric execution. Repo
  • Nextflow IaC Plugin: Infrastructure-as-code orchestration for pipelines across Arm, GPU, and x86 infrastructure. Repo
  • Distributed MCP Protocol: Distributed protocol design for AI-native content delivery and agent coordination. Demo
  • PCIe-Net and RDMA over PCIe/CXL: TCP/IP-over-PCIe/CXL and RDMA-style data movement across high-speed interconnect fabrics. Demo
  • CXL Switch SoC and Cluster-on-Board: Switch-chip and multi-CPU board-level architectures for scalable AI and HPC systems.
  • Multi-rail HPC Computing System: Multi-rail HPC computing architecture for production rendering workloads.

Education and affiliations

Background

  • Ph.D., University of Wisconsin - Madison
  • M.S., University of Wisconsin - Madison
  • B.S., University of Science and Technology of China
  • Voting member of Linux Foundation Edge and Akraino Project, 2022 and 2023
  • Vice Director of the Film Advanced Technology Committee of CSMPTE, 2018
  • Industry Professorship of Jiangnan University, 2017-2021

Patents

Selected patent record

  • US 20120290696: Longest Prefix Matching of Variable-Sized Hierarchical Names by Treelets
  • CN114827151A: Heterogeneous clustered devices and servers based on PCIe, CXL, and UCIe physical links
  • CN114745325A: MAC in MAC network encoding based on PCIe, CXL, and UCIe physical links
  • CN110891081A: Packet sending, routing, broadcasting, and receiving
  • CN111209098A: Intelligent rendering scheduling
  • CN107819704A: Scalable wireless media application for edge computing

Contact

Follow the work.

I keep the site focused on public technical writing, project pages, and systems experiments.