LLM Routing, Research, and AI Infra.
Papers and infrastructure work around routing as a systems problem.
Infrastructure and standards work beyond the router.
Maintainer, steering, and reviewer roles across gateways, service mesh, and inference infrastructure.
vLLM Semantic Router
Co-Founder
An open semantic decision layer for preference-aligned Mixture-of-Models.
Elephant Agent
Creator
Personal-model-first self-evolving AI agent that grows correctable understanding and gets curious at the user's pace.
Inferoa
Builder
Inference-native tokenmaxxing agent harness for loop engineering.
Envoy Gateway
Steering Committee and Maintainer
Manages Envoy Proxy as a standalone or Kubernetes-based application gateway.
Envoy AI Gateway
Maintainer
Manages unified access to generative AI services built on Envoy Gateway.
vLLM AIBrix
Maintainer
Cost-efficient and pluggable infrastructure components for GenAI inference.
Higress
Approver
AI gateway and AI-native API gateway.
Istio
Maintainer
Connects, secures, controls, and observes services.
Kiali
Maintainer
Observability console for Istio with service mesh.
Aeraki Mesh
Maintainer
Manages any layer-7 protocols in a service mesh.
Merbridge
Maintainer
Uses eBPF to speed up service mesh data paths.
Kubernetes Gateway API
Reviewer
Role-oriented, portable, and expressive interfaces for Kubernetes networking.
Kubernetes Ingress2Gateway
Reviewer
Converts Ingress resources to Gateway API resources.
Papers on routing, systems, and inference optimization.
A research archive spanning semantic routing, agent behavior, and infrastructure efficiency.
Fast and Faithful: Real-Time Verification for Long-Document Retrieval-Augmented Generation Systems
SIGIR 2026 Industry Track
Presents a real-time verification layer for long-document RAG systems, handling contexts up to 32K tokens while balancing latency with grounding coverage for interactive production deployments.
Elephant Agent: Personal-Model-First Self-Evolution for Personal AI
Agentic Intelligence Lab
Introduces a personal-model-first agent architecture where personal AI grows a correctable understanding of Identity, World, Pulse, and Journey through user-paced curiosity and reflection after each turn.
Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference
arXiv Technical Report
Proposes token-budget-aware pool routing that estimates each request's token budget online and dispatches it to short- or long-context serving pools, reducing GPU cost while improving stability for LLM inference.
vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models
arXiv Technical Report
Introduces vLLM Semantic Router as a signal-driven routing framework for mixture-of-modality deployments, composing heterogeneous signals into deployment-specific policies across cost, privacy, latency, and safety.
The Workload-Router-Pool Architecture for LLM Inference Optimization: A Vision Paper from the vLLM Semantic Router Project
arXiv Technical Report
Synthesizes recent routing, fleet, multimodal, and governance work into the Workload-Router-Pool architecture, framing routing as one layer in a broader inference optimization stack.
Visual Confused Deputy: Exploiting and Defending Perception Failures in Computer-Using Agents
arXiv Technical Report
Formalizes the visual confused deputy as a security failure mode for computer-use agents and proposes a dual-channel guardrail for validating click targets and action reasoning before execution.
Outcome-Aware Tool Selection for Semantic Routers: Latency-Constrained Learning Without LLM Inference
arXiv Technical Report
Presents OATS, an offline embedding-refinement method that improves semantic-router tool ranking under single-digit millisecond CPU budgets without adding serving-time model inference.
Adaptive Vision-Language Model Routing for Computer Use Agents
arXiv Technical Report
Proposes Adaptive VLM Routing to estimate step difficulty in computer-use agents and route each action to the cheapest model that can still satisfy a target reliability threshold.
98x Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router
arXiv Technical Report
Combines flash attention, prompt compression, and near-streaming body processing to cut routing latency from seconds to tens of milliseconds on lightweight shared serving hardware.
inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
arXiv Technical Report
Introduces a queueing-theory-grounded fleet planner and discrete-event simulator for sizing multi-pool LLM GPU fleets against P99 TTFT targets without requiring up-front hardware profiling.
FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
arXiv Technical Report
Derives the minimum-cost two-pool LLM fleet directly from workload distributions and P99 TTFT targets, then maps the optimal boundary to a deployable compress-and-route strategy.
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
arXiv Technical Report
Derives the 1/W law, showing that tokens per watt roughly halve whenever the serving context window doubles, making context-length routing a first-order energy-efficiency lever.
Conflict-Free Policy Languages for Probabilistic ML Predicates: A Framework and Case Study with the Semantic Router DSL
arXiv Technical Report
Shows how probabilistic ML predicates in policy languages can silently co-fire on the same query and adds conflict detection plus a softmax prevention mechanism in the Semantic Router DSL.
From Inference Routing to Agent Orchestration: Declarative Policy Compilation with Cross-Layer Verification
arXiv Technical Report
Extends the Semantic Router DSL from stateless per-request routing to multi-step agent workflows, compiling verified decision artifacts across orchestration, Kubernetes, and protocol layers.
Knowledge Access Beats Model Size: Memory Augmented Routing for Persistent AI Agents
arXiv Technical Report
Shows that conversational memory and retrieval-grounded routing let a lightweight 8B model recover most of a much larger model's performance on persistent user-specific queries while dramatically reducing cost.
When to Reason: Semantic Router for vLLM
NeurIPS - MLForSys
Routes prompts by reasoning requirements so reasoning is only invoked when it pays off, improving accuracy while cutting token usage and latency versus always-on reasoning.
Category-Aware Semantic Caching for Heterogeneous LLM Workloads
arXiv Technical Report
Proposes category-aware semantic caching with category-specific similarity thresholds, TTLs, and quotas, using a hybrid split between in-memory HNSW retrieval and external document storage.
Community responsibilities across open source infrastructure.
Current and past roles that connect research, implementation, review, and standards work.
Agentic Intelligence Lab
Chair
Chairing the lab's research and community work on agentic AI, personal AI agents, and system intelligence.
Kubernetes AI Gateway WorkGroup
Co-Chair
Leading the community effort to define standards for AI Gateway in the Kubernetes ecosystem.
CNCF Ambassador
Fall 2023 Ambassador
Representing and promoting Cloud Native Computing Foundation projects and values globally.
Linux Foundation APAC Open Source Evangelist
2024 Program
Advocating for open source adoption and best practices across the Asia-Pacific region.
KubeCon Program Committee
KubeCon 2024 Hong Kong
Reviewing and selecting talks for one of the largest cloud-native conferences.