Helium AI Router
High-throughput OpenAI-compatible LLM gateway, MCP SSE server, and multi-model consensus routing engine.
Overview
Helium (router.y7.hk) is an enterprise-grade LLM API gateway and orchestration engine built to provide unified, resilient access to 40+ frontier AI models. Designed as a drop-in OpenAI-compatible proxy (/v1) and Model Context Protocol (MCP) SSE host, Helium unifies routing across Google Antigravity (Gemini 3.8 & Claude 4.6), Xiaomi MiMo, DeepSeek V4, Alibaba Qwen, and xAI Grok under a single authenticated interface.
Key Architecture & Features
- Antigravity Protocol Alignment: Deep integration with Google Cloud Code internal APIs featuring singleton account pooling, thinking-budget auto-rescaling, and continuous 3-second SSE heartbeats.
- Thinking Dropout Auto-Healing: Proprietary stream interception that detects empty reasoning terminations (
STOPwith 0 content tokens) and auto-completes requests seamlessly in the same connection without client drops. - Wave Multi-Model Consensus Engine: Orchestrates multi-expert parallel reasoning (up to 4 tiers including multi-stage pipelines) synthesized in real time into authoritative responses.
- Native MCP Server: Full SSE transport implementation exposing database metrics, latency benchmarking, and automated key lifecycle tools directly to AI coding agents.
- High-Velocity Telemetry: Nothing Phone-inspired telemetry dashboard built with Next.js 16 and GSAP, visualizing 10-second rolling token speed, live in-flight sessions, and token routing Sankey flows.
Results & Impact
Helium acts as the central intelligence nervous system across all internal AI agents and coding tools, processing millions of tokens daily with sub-second failover recovery, automatic quota defense, and zero-downtime hot reloading.