The next era of computing will not be organized around one model. It will be a changing pool of large and small, open and closed, local and cloud intelligence.
My work is about making that diversity usable as one system. A user should choose an objective—speed, cost, balance, or accuracy—while semantic routing decides the capability path across models, tools, memory, verifiers, hardware, and locations.
- Preference-aligned Mixture-of-Models gives applications one stable model contract while the system handles execution.
- Semantic routing turns workload signals and policy into inspectable decisions across a changing model pool.
- Workload–Router–Pool co-design connects semantic demand with live serving state, cache, placement, and heterogeneous hardware.
I pursue this work in the open through vLLM Semantic Router, research, and collaboration across the inference ecosystem.