As of 2026-08 , AlphaMoat tracks 330 AI websites tagged AI Observability & Eval . The most visited is Arena with 29.9M monthly visits (-7% MoM). Ranked by latest monthly traffic, updated monthly.
# Product Company Industry Visits MoM
Public LLM leaderboard built from community votes
LMArena News & Analytics 29.9M -7% Independently compare AI model and API-provider performance.
Artificial Analysis General AI 7.7M +25.8% AI observability and security platform for any tech stack.
Datadog Programming 5.7M -6.2% Automatically diagnose product issues, fix bugs, and generate code changes.
PostHog Programming 3.4M +3.5% Help development teams find errors and monitor application performance.
Sentry Programming 3.0M -8.4% A training-data platform combining human expertise and technology for agents and LLMs.
Toloka Programming 2.4M -9.2% Track, visualize, and improve machine-learning experiment workflows
Weights & Biases General AI 2.0M -4.6% An arena where user votes seat design models.
Design Arena General AI 1.3M -11.3% An independent AI model leaderboard comparing GPT, Claude, Gemini, Llama and more on a single page.
LLM Stats General AI 1.2M +5.6% A cloud platform unifying metrics, logs, and traces with troubleshooting aid.
Grafana Programming 1.1M -9.6% A monitoring and community-info platform for OpenAI Codex users.
CodexRadar General AI 946.3K +3.2% An AI DevOps platform spanning delivery, testing, security, and cost optimization.
Harness Programming 871.9K -2.2% A production-ready OpenAI-compatible AI gateway that unifies multi-model routing.
OrcaRouter General AI 833.0K +931.4% Open-source tracing, evaluation, and continuous improvement for production AI agents.
Langfuse General AI 828.9K -13.4% Researching the capabilities of intelligent compute models and industry development trends.
Epoch AI News & Analytics 608.8K +12.3% Provides training data, evaluation, and full-stack tech to build reliable intelligent systems
Scale General AI 569.9K -4.3% An AI SRE unifying incident response, on-call, and log-metric tracking.
Better Stack Programming 550.0K +2.7% End-to-end AI engineering stack for production-grade agents.
Pydantic Programming 454.4K -15.3% Independently assessing frontier AI capability risk and real-work impact.
METR General AI 446.7K +76.2% Track, evaluate, debug and improve AI apps and agents
Comet ML General AI 282.8K -30.7% Incubate open, accessible, and scalable large-model systems.
LMSYS General AI 267.8K -7.4% Unified telemetry and fast observable queries for AI-era software.
Hound Technology General AI 240.4K -19.1% Track evaluations in real time and locate AI agent regression issues.
Braintrust General AI 238.5K -16.2% Quantify individuals' and teams' AI coding activity and cost
WakaTime Programming 236.0K -20.2% Continuously track, evaluate and experiment to improve AI agents in production.
Arize AI General AI 235.4K -8% Provide gateway, routing, logging and prompt management for generative AI apps.
Portkey General AI 226.2K -15% A platform that gathers real human feedback and continuously evaluates voice and conversation AI.
Hume AI Audio 211.3K -20.4% Open-source platform to manage, evaluate, and monitor agents, LLMs, and ML.
MLflow General AI 186.2K -19.2% Open-source labeling for multimodal training data, agent traces, and LLM evaluation.
HumanSignal General AI 185.6K -3.7% Data development and evaluation infrastructure for frontier LLMs.
Snorkel AI General AI 185.1K -6.9% Discover real-time attack signals across networks, identity, cloud, and AI environments
Vectra Programming 175.7K +1.7% Independent AI evaluation benchmark platform
Vals General AI 167.3K -11.3% Simulate, evaluate and observe generative AI agents in real time
H3 Labs General AI 162.5K +38.6% A tool that sends one prompt to 25 models for consensus comparison.
PromptQuorum General AI 161.2K +20.2% Open-source evaluation and red-teaming for LLM apps and agents
Promptfoo General AI 158.2K -13.7% Open-source evaluation and monitoring for LLM retrieval and machine learning systems
Evidently AI General AI 157.2K -5.8% Simulate, evaluate, monitor and improve AI agents in production
Future AGI General AI 117.0K +43.3% Unifies LLM evaluation, tracing, and production observability for enterprises.
Confident AI General AI 114.8K -1.3% A multimodal data curation platform for physical AI
Voxel51 General AI 108.3K -17.7% Observe, evaluate, and manage production AI models and prompts.
Keywords AI General AI 89.1K +28.8% An LLMOps platform that routes and monitors large model requests in one place
Helicone General AI 85.2K -17.3% EU-hosted unified AI agent gateway across models
Opper Programming 82.1K +5.2% Enterprise-grade AI data and implementation services covering multimodal collection, labeling, and evaluation.
YPAI General AI 79.8K +7.2% A free AI paper-review tool from Stanford.
Stanford Agentic Reviewer Education & Research 71.2K -42.8% Sovereign-grade AI gateway and agent platform.
Orq.ai General AI 70.0K +0.8% An automated testing platform for voice and conversation agents.
Cekura General AI 66.2K +35.5% An AI model benchmark on real coding tasks.
World of AI Bench General AI 64.0K +50.8% A visual large-model arena to compare generation effects side by side.
MagicArena General AI 60.5K -28.8% Evaluate and monitor performance risks in machine learning models
Fiddler AI General AI 57.2K +3.4% An intelligence platform measuring AI engineering spend, output quality, and return.
Milestone Programming 52.6K -4.2% Previous 1 2 3 … 7 Next Data month: 2026-08 · Visit figures are third-party traffic estimates aggregated monthly — best for scale and trends. Multiple domains of the same product are merged. See methodology .