Full-stack AI infrastructure

Full-stack AI infrastructure for GPUs, clouds, models, and agents.

I work across the full stack of AI infrastructure: GPU systems and clusters, high-performance networking, cloud runtimes, model and memory systems, and agent engineering. This site is a compact index of my current writing and projects.

01 GPU systems

GPU clusters, RDMA/RoCE, NVLink, PCIe/CXL, memory pressure, and model execution.

02 Cloud runtime

Kubernetes, OKE, FSS, heterogeneous compute, workflow placement, and cost-performance.

03 Model systems

MoE inference, vLLM, llama.cpp, KV cache behavior, GPU prefetch, and CPU memory banks.

04 Agent systems

Agent control planes, MCP, durable memory, workspaces, deterministic workflows, and review loops.

Work and interests

AI Infrastructure Architect with 15+ years across high-performance computing, distributed systems, advanced networking, and large-scale AI infrastructure.

My current focus is full-stack AI infrastructure: how accelerators, fabrics, memory, cloud runtimes, model systems, and agent systems fit together as AI software becomes more persistent, distributed, and infrastructure-aware.

  • GPU and fabricCluster networking, remote PCIe/CXL, RDMA-style movement, and accelerator composition.
  • Model runtimeMemory-efficient MoE inference, GPU prefetch, CPU offload, and local inference constraints.
  • Cloud architectureOKE/FSS, heterogeneous Nextflow workflows, cost-performance, and persistent runtime foundations.
  • Agent engineeringAgent Graph, Agent Runner, Memory Agent, MCP control planes, and deterministic human-agent workflows.

Recent writing

A short list first. Expand for the full archive.

Show full blog list

Recent projects

Recent agent and infrastructure projects first. Expand for the full project list.

  • Agent Graphagent control plane

    Deterministic multi-agent engineering for VS Code: issue graphs, MCP control, isolated Codex workers, and parallel execution through bounded workflows.

    Project pageInstall VSIX
  • Agent Runneragent runtime

    Server-side multi-agent orchestration with mailbox queues, agenda tasks, groups, resources, and deterministic project state.

    Project page
  • Agentify Cloudcloud runtime

    Agent-native cloud runtime with FastAPI, FastMCP, AGENTS.md policy, route intents, JSON contracts, and validation.

    Project page
  • vLLM-MoE GPU Prefetchmodel runtime

    Reduces a normal 80GB GPU demand to 40GB while reaching 128.59 tokens/s on an A100 40GB for 26B MoE inference.

    Repo
  • PCIe over Networkremote fabric

    Software-defined remote PCIe virtualization where a host can access physical PCIe devices across a network while preserving native driver behavior.

    BlogDemo
Show full project list
  • vLLM MoE CPU Offloadmodel runtime

    GPU-native MoE offload using CPU host memory as the expert-weight bank for constrained GPU memory.

    Repo
  • llama.cpp-MoElocal inference

    Router-aware GPU expert slots for local MoE inference under constrained GPU memory.

    Repo
  • GPU-native Scheduler for GPU Computingscheduling

    Patent-submitted scheduling approach for GPU resource management.

  • SuperKernel for SuperPod GPU Clustersgpu fabric

    Patent-submitted Jupyter kernel architecture for large GPU fabric execution.

    Repo
  • Nextflow IaC Pluginheterogeneous cloud

    Infrastructure-as-code orchestration for pipelines across Arm, GPU, and x86 infrastructure.

    RepoReference 1Reference 2
  • Distributed MCP Protocol for AI-native CDN Architecturedistributed agents

    Distributed protocol design for AI-native content delivery and agent coordination.

    Demo
  • PCIe-Net and RDMA over PCIe/CXLfabric networking

    TCP/IP-over-PCIe/CXL and RDMA-style data movement across high-speed interconnect fabrics.

    Demo
  • CXL Switch SoC and Cluster-on-Boardsystem architecture

    Switch-chip and multi-CPU board-level architectures for scalable AI and HPC systems.

  • Multi-rail HPC Computing Systemhpc rendering

    Multi-rail HPC computing architecture for production rendering workloads.

Follow the work

Public writing and project updates are kept here, on LinkedIn, and on GitHub.