Full-stack AI infrastructure
Full-stack AI infrastructure for GPUs, clouds, models, and agents.
I work across the full stack of AI infrastructure: GPU systems and clusters, high-performance networking, cloud runtimes, model and memory systems, and agent engineering. This site is a compact index of my current writing and projects.
GPU clusters, RDMA/RoCE, NVLink, PCIe/CXL, memory pressure, and model execution.
Kubernetes, OKE, FSS, heterogeneous compute, workflow placement, and cost-performance.
MoE inference, vLLM, llama.cpp, KV cache behavior, GPU prefetch, and CPU memory banks.
Agent control planes, MCP, durable memory, workspaces, deterministic workflows, and review loops.
Summary
Work and interests
AI Infrastructure Architect with 15+ years across high-performance computing, distributed systems, advanced networking, and large-scale AI infrastructure.
My current focus is full-stack AI infrastructure: how accelerators, fabrics, memory, cloud runtimes, model systems, and agent systems fit together as AI software becomes more persistent, distributed, and infrastructure-aware.
- GPU and fabricCluster networking, remote PCIe/CXL, RDMA-style movement, and accelerator composition.
- Model runtimeMemory-efficient MoE inference, GPU prefetch, CPU offload, and local inference constraints.
- Cloud architectureOKE/FSS, heterogeneous Nextflow workflows, cost-performance, and persistent runtime foundations.
- Agent engineeringAgent Graph, Agent Runner, Memory Agent, MCP control planes, and deterministic human-agent workflows.
Recent blogs
Recent writing
A short list first. Expand for the full archive.
- Scaling 1,000 AI Agents on OCI Kubernetes Engine and File Storage
Agent fleet architecture using OKE, persistent file systems, and governed runtime foundations.
- Agent Graph 0.2.0: Parallel Issues, One Deterministic Control Plane
Parallel agent execution with observable issue graphs, leases, and bounded control.
- Agent Graph: A Deterministic Control Plane for Multi-Agent Engineering
A project architecture discussion for software that contains agents inside the workflow.
- PCIe over the Network: When a Remote GPU Looks Local
Remote PCIe virtualization for composing accelerator resources across system boundaries.
- Raising an Agent: From Execution to Self-Evolution
A four-level view of agent capability, from orchestration to self-evolving methods.
Show full blog list
- The Contextual Turn: How AI Is Changing Our Understanding of Language
Language, context, memory, and interaction in AI systems.
- Agentify Cloud: When the Agent Becomes the Cloud Runtime
Agent-native cloud runtime design with policy, route intents, contracts, and validation.
- What Is Tau Computing?
Compute, memory, and fabric design for future AI systems.
- Zettascale in Practice: MRC Benefit
Systems-level thinking for AI networking and infrastructure efficiency.
- MRC and the Future of AI Networking
Network architecture for AI-scale communication and resource composition.
- Run Nextflow with heterogeneous computing on OCI
Pipeline-driven heterogeneous computing across Arm, GPU, and x86 cloud resources.
- Cut Nextflow costs by 70% with OCI
Cost-aware workflow placement for heterogeneous pipelines.
- Zettascale in practice: Scaling beyond limits
Infrastructure notes on scaling behavior, benchmarks, and large AI workload limits.
- Zettascale OSU and NCCL benchmark for H100 AI workloads
Benchmarking RDMA and GPU cluster communication for H100 AI workloads.
- The Story of OpenClaw: Learning to Collaborate
Human-guided AI engineering and system building.
- Human in the Loop: Guiding AI to Break the Bus/Network Architecture Barrier
Guided AI exploration for bus, network, and architecture design.
- Human-in-the-Loop Engineering and Vibe Coding
Human direction, agent execution, and engineering review.
- AI Coding BucketFS: A Transactional FUSE Filesystem for Object Storage
A storage and filesystem experiment shaped by AI-assisted implementation.
- AI Networking: How Fractal Scales Beyond Limits
AI networking notes on scaling, limits, and architecture tradeoffs.
- Inside AI Infrastructure, Series II: Benchmarking
Benchmarking as a way to reason about infrastructure behavior under workload pressure.
- A New Architecture Shift: NVIDIA, Enfabrica, and Intel
Accelerator fabrics and the changing shape of AI infrastructure.
Recent projects
Recent projects
Recent agent and infrastructure projects first. Expand for the full project list.
- Agent Graph
Deterministic multi-agent engineering for VS Code: issue graphs, MCP control, isolated Codex workers, and parallel execution through bounded workflows.
Project pageInstall VSIX - Agent Runner
Server-side multi-agent orchestration with mailbox queues, agenda tasks, groups, resources, and deterministic project state.
Project page - Agentify Cloud
Agent-native cloud runtime with FastAPI, FastMCP, AGENTS.md policy, route intents, JSON contracts, and validation.
Project page - vLLM-MoE GPU Prefetch
Reduces a normal 80GB GPU demand to 40GB while reaching 128.59 tokens/s on an A100 40GB for 26B MoE inference.
Repo
Show full project list
- vLLM MoE CPU Offload
GPU-native MoE offload using CPU host memory as the expert-weight bank for constrained GPU memory.
Repo - llama.cpp-MoE
Router-aware GPU expert slots for local MoE inference under constrained GPU memory.
Repo - GPU-native Scheduler for GPU Computing
Patent-submitted scheduling approach for GPU resource management.
- SuperKernel for SuperPod GPU Clusters
Patent-submitted Jupyter kernel architecture for large GPU fabric execution.
Repo - Nextflow IaC Plugin
Infrastructure-as-code orchestration for pipelines across Arm, GPU, and x86 infrastructure.
RepoReference 1Reference 2 - Distributed MCP Protocol for AI-native CDN Architecture
Distributed protocol design for AI-native content delivery and agent coordination.
Demo - PCIe-Net and RDMA over PCIe/CXL
TCP/IP-over-PCIe/CXL and RDMA-style data movement across high-speed interconnect fabrics.
Demo - CXL Switch SoC and Cluster-on-Board
Switch-chip and multi-CPU board-level architectures for scalable AI and HPC systems.
- Multi-rail HPC Computing System
Multi-rail HPC computing architecture for production rendering workloads.
Contact
Follow the work
Public writing and project updates are kept here, on LinkedIn, and on GitHub.