This project is no longer maintained. Some features may be non-functional; see Funnel instead.
Multi-LLM evaluation platform querying 6+ models simultaneously with synthesized consensus responses and support for 11+ APIs through a unified routing layer.
Overview
Consensus was a sandbox web app built to test multiple LLMs side-by-side in real time. It ran prompts across 6+ Gemini models at the same time, displayed streaming responses in parallel, and combined their answers to spot where models agreed or disagreed.
Highlights
Streamed responses at 236 tokens/s with 2.5s p95 latency using Gemini 2.0 Flash.
Ran 6+ models concurrently with side-by-side LaTeX and code rendering in a collapsible layout.
Used a React and Express backend to parse and merge outputs across multiple models.

