Technical notes on AI systems
Building software where agents are inside the system, not just beside it.
My current work sits between GPU infrastructure, cloud runtimes, and agent engineering. The question I keep returning to is simple: what changes when an application is designed with agents as durable parts of the architecture?
Work and interests
Less resume, more technical workbench.
AI Infrastructure Architect with 15+ years across high-performance computing, distributed systems, advanced networking, and large-scale AI infrastructure.
I use this site as a technical notebook for systems that move AI from model execution into persistent software: GPU clusters, remote PCIe/CXL fabrics, Kubernetes agent fleets, memory systems, and human-guided agent workflows.
The professional scope is intentionally layered. The lower layer asks how compute and memory move. The middle layer asks how agents run with state, identity, and audit. The upper layer asks how humans and agents collaborate on real engineering work.
Three layers
A practical map for agent-native software.
Use the tabs to move between the layers. This is the main narrative of the site: infrastructure first, governed agent runtime second, experimentation third.
GPU systems and clusters
Cloud GPU systems, RDMA networking, NVLink environments, disaggregated PCIe/CXL infrastructure, memory pooling, and accelerator composition. This layer is about the physical and logical substrate where large AI workloads run.
Recent writing
Architecture notes, not marketing posts.
The writing is grouped by the same layers: GPU systems, agent-native infrastructure, agent engineering, and cloud architecture.
-
Scaling 1,000 AI Agents on OCI Kubernetes Engine and File Storage
A practical agent fleet architecture using OKE, persistent file systems, and governed runtime foundations.
-
Agent Graph 0.2.0: Parallel Issues, One Deterministic Control Plane
How parallel agent execution can stay observable through issue graphs, leases, and bounded control.
-
Agent Graph: A Deterministic Control Plane for Multi-Agent Engineering
A design discussion for moving from assistant-driven coding to software with agents inside the workflow.
-
PCIe over the Network: When a Remote GPU Looks Local
Remote PCIe virtualization as a way to compose accelerator resources across system boundaries.
-
Raising an Agent: From Execution to Self-Evolution
A four-level view of agent capability, from orchestration to self-evolving methods.
-
The Contextual Turn: How AI Is Changing Our Understanding of Language
A language and context note connecting AI systems to meaning, memory, and interaction.
-
Agentify Cloud: When the Agent Becomes the Cloud Runtime
A proposal for agent-native cloud runtime design where policy, routes, and validation shape execution.
-
What Is Tau Computing?
An architectural concept note on compute, memory, and fabric design for future AI systems.
-
Zettascale in Practice: MRC Benefit
Systems-level thinking for large-scale AI networking and infrastructure efficiency.
-
MRC and the Future of AI Networking
Network architecture notes for AI-scale communication and memory/resource composition.
-
Run Nextflow with heterogeneous computing on OCI
Pipeline-driven heterogeneous computing across Arm, GPU, and x86 cloud resources.
-
Cut Nextflow costs by 70% with OCI
Cost-aware workflow placement for heterogeneous scientific and AI pipelines.
-
Zettascale in practice: Scaling beyond limits
Infrastructure notes on scaling behavior, benchmarks, and large AI workload limits.
-
Zettascale OSU and NCCL benchmark for H100 AI workloads
Benchmarking RDMA and GPU cluster communication for H100 AI workloads.
-
The Story of OpenClaw: Learning to Collaborate
A collaboration note on human-guided AI engineering and system building.
-
Human in the Loop: Guiding AI to Break the Bus/Network Architecture Barrier
How guided AI exploration connects to bus, network, and architecture design.
-
Human-in-the-Loop Engineering and Vibe Coding
A practical framing of human direction, agent execution, and engineering review.
-
AI Coding BucketFS: A Transactional FUSE Filesystem for Object Storage
A storage and filesystem experiment shaped by AI-assisted implementation.
-
AI Networking: How Fractal Scales Beyond Limits
AI networking notes on scaling, limits, and architecture tradeoffs.
-
Inside AI Infrastructure, Series II: Benchmarking
Benchmarking as a way to reason about infrastructure behavior under real workload pressure.
-
A New Architecture Shift: NVIDIA, Enfabrica, and Intel
A systems architecture note on accelerator fabrics and the changing shape of AI infrastructure.
Selected projects
Projects as evidence for the writing.
The homepage keeps projects concise. Each one points to a system idea: control planes, memory pressure, remote fabrics, or heterogeneous execution.
-
Agent Graph
Deterministic multi-agent engineering for VS Code: issue graphs, MCP control, isolated Codex workers, and parallel execution through bounded workflows.
Project page · Install VSIX -
Agent Runner
Server-side multi-agent orchestration with mailbox queues, agenda tasks, groups, resources, and deterministic project state.
Project page -
Agentify Cloud
Agent-native cloud runtime with FastAPI, FastMCP, AGENTS.md policy, route intents, JSON contracts, and validation.
Project page -
vLLM-MoE GPU Prefetch
Reduces a normal 80GB GPU demand to 40GB while reaching 128.59 tokens/s on an A100 40GB for 26B MoE inference.
Repo -
PCIe over Network
Software-defined remote PCIe virtualization where a host can access physical PCIe devices across a network while preserving native driver behavior.
Blog · Demo
Show more projects
- vLLM MoE CPU Offload: GPU-native MoE offload using CPU host memory as the expert-weight bank. Repo
- llama.cpp-MoE: Router-aware GPU expert slots for local MoE inference under constrained GPU memory. Repo
- GPU-native Scheduler for GPU Computing: Patent-submitted scheduling approach for GPU resource management.
- SuperKernel for SuperPod GPU Clusters: Patent-submitted Jupyter kernel architecture for large GPU fabric execution. Repo
- Nextflow IaC Plugin: Infrastructure-as-code orchestration for pipelines across Arm, GPU, and x86 infrastructure. Repo
- Distributed MCP Protocol: Distributed protocol design for AI-native content delivery and agent coordination. Demo
- PCIe-Net and RDMA over PCIe/CXL: TCP/IP-over-PCIe/CXL and RDMA-style data movement across high-speed interconnect fabrics. Demo
- CXL Switch SoC and Cluster-on-Board: Switch-chip and multi-CPU board-level architectures for scalable AI and HPC systems.
- Multi-rail HPC Computing System: Multi-rail HPC computing architecture for production rendering workloads.
Education and affiliations
Background
- Ph.D., University of Wisconsin - Madison
- M.S., University of Wisconsin - Madison
- B.S., University of Science and Technology of China
- Voting member of Linux Foundation Edge and Akraino Project, 2022 and 2023
- Vice Director of the Film Advanced Technology Committee of CSMPTE, 2018
- Industry Professorship of Jiangnan University, 2017-2021
Patents
Selected patent record
- US 20120290696: Longest Prefix Matching of Variable-Sized Hierarchical Names by Treelets
- CN114827151A: Heterogeneous clustered devices and servers based on PCIe, CXL, and UCIe physical links
- CN114745325A: MAC in MAC network encoding based on PCIe, CXL, and UCIe physical links
- CN110891081A: Packet sending, routing, broadcasting, and receiving
- CN111209098A: Intelligent rendering scheduling
- CN107819704A: Scalable wireless media application for edge computing
Contact
Follow the work.
I keep the site focused on public technical writing, project pages, and systems experiments.