InfrastructureAI GatewayGoRustLLM RoutingMCPNext.js

Helium AI Router

High-throughput OpenAI-compatible LLM gateway, MCP SSE server, and multi-model consensus routing engine.

40+
Supported Models
1M Tokens
Max Context
< 1.2s
Edge TTFT

Overview

Helium (router.y7.hk) is an enterprise-grade LLM API gateway and orchestration engine built to provide unified, resilient access to 40+ frontier AI models. Designed as a drop-in OpenAI-compatible proxy (/v1) and Model Context Protocol (MCP) SSE host, Helium unifies routing across Google Antigravity (Gemini 3.8 & Claude 4.6), Xiaomi MiMo, DeepSeek V4, Alibaba Qwen, and xAI Grok under a single authenticated interface.

Key Architecture & Features

  • Antigravity Protocol Alignment: Deep integration with Google Cloud Code internal APIs featuring singleton account pooling, thinking-budget auto-rescaling, and continuous 3-second SSE heartbeats.
  • Thinking Dropout Auto-Healing: Proprietary stream interception that detects empty reasoning terminations (STOP with 0 content tokens) and auto-completes requests seamlessly in the same connection without client drops.
  • Wave Multi-Model Consensus Engine: Orchestrates multi-expert parallel reasoning (up to 4 tiers including multi-stage pipelines) synthesized in real time into authoritative responses.
  • Native MCP Server: Full SSE transport implementation exposing database metrics, latency benchmarking, and automated key lifecycle tools directly to AI coding agents.
  • High-Velocity Telemetry: Nothing Phone-inspired telemetry dashboard built with Next.js 16 and GSAP, visualizing 10-second rolling token speed, live in-flight sessions, and token routing Sankey flows.

Results & Impact

Helium acts as the central intelligence nervous system across all internal AI agents and coding tools, processing millions of tokens daily with sub-second failover recovery, automatic quota defense, and zero-downtime hot reloading.