Data month 2026-07

Best AI Observability & Eval Tools

As of 2026-07, AlphaMoat tracks 253 AI websites tagged AI Observability & Eval. The most visited is LMArena with 31.8M monthly visits (-0.6% MoM). Ranked by latest monthly traffic, updated monthly.

#ProductIndustryVisitsMoM
1
LMArena
LMArena
Prompt once. Compare multiple AI-built apps for free.
Business31.8M-0.6%
2
Artificial Analysis
Artificial Analysis
AI Model & API Providers Analysis 是 Artificial Analysis 旗下与 artificialanalysis.ai 对应的产品官网,主要提供:Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency.
Programming6.1M+39.6%
3
Datadog
Datadog
Datadog 对应 datadoghq.com。See inside any stack, any app, at any scale, anywhere.
Programming6.1M+2.7%
4
Sentry
Sentry
Sentry 对应 sentry.io。Application performance monitoring for developers & software teams to see errors clearer, solve issues faster & continue learning continuously. Get started at sentry.io.
Education3.3M+4.9%
5
PostHog
PostHog
PostHog – We make dev tools for product engineers 对应 posthog.com。All your developer tools in one place. PostHog gives engineers everything to build, test, measure, and ship successful products faster. Get started free.
Programming3.2M+12.9%
6Programming2.1M+1.9%
7
Design Arena
Design Arena
基于真实用户投票的 AI 设计能力评测平台。
General AI1.4M-0.1%
8
Grafana Labs
Grafana Labs
Grafana Labs 对应 grafana.com。Avoid lock-in and ensure reliability with Grafana Cloud. The AI-powered platform built on open standards unifies metrics, logs, traces, profiles, and business data.
Business1.3M-5.4%
9
LLM Stats
LLM Stats
Compare API models by benchmarks, cost & capabilities
Programming1.1M+15.5%
10
Langfuse
Langfuse
开源LLM应用程序分析
Programming956.6K-0.1%
11
CodexRadar
CodexRadar
Codex 今天降智了吗?
General AI916.5K+341.5%
12
Scale
Scale
Scale AI的核心功能包括高质量的训练数据,由经验丰富的专家团队进行的数据标记和注释,用户友好的平台界面以及可满足各种AI应用程序需求的可扩展性。
Image595.6K-1.7%
13
Epoch AI
Epoch AI
Epoch AI 是 Epoch AI 旗下与 epoch.ai 对应的产品官网,主要提供:Epoch AI is a research institute investigating key trends and questions that will shape the trajectory and governance of Artificial Intelligence.
General AI541.9K+23.8%
14
Better Stack
Better Stack
Better Stack 是 Better Stack 旗下与 betterstack.com 对应的产品官网,主要提供:AI SRE and MCP server, incident management, on-call, logs, metrics, traces, and error tracking. 7,000+ happy customers. 60-day money back guarantee.
General AI535.5K-12.7%
15
Arena
Arena
Arena 对应 lmarena.ai。Chat, compare, vote for the world's best AI models. Join the community shaping the public leaderboard for LLMs, image, and code models through real-world evaluation.
AI Marketing414.7K-3.4%
16
Comet
Comet
Comet is represented by comet.com. Comet provides an end-to-end model evaluation platform for AI developers, with best-in-class LLM evaluations, experiment tracking, and production monitoring.
Programming408.0K-22.5%
17
Honeycomb
Honeycomb
Honeycomb 对应 honeycomb.io。Honeycomb is the observability platform built for AI-era software. Fast queries, unified telemetry, and LLM observability. Used by Slack, Intercom, and Dropbox.
Programming297.1K+7.7%
18
WakaTime
WakaTime
开发人员编程统计和分析仪表板
Programming295.6K-6.2%
19
lmsys
lmsys
开源聊天机器人,性能接近ChatGPT
General AI289.2K+6.1%
20
Braintrust
Braintrust
Braintrust 对应 braintrust.dev。Braintrust - The AI observability platform for building quality AI products
Programming284.8K+20.7%
21
Portkey
Portkey
可观察性套件AI网关提示播放器
General AI266.1K+8%
22
Hume AI
Hume AI
移情语音接口 (EVI) 表情测量API自定义模型API
Audio265.4K+7.4%
23
Arize AI
Arize AI
监控仪表板评估和性能跟踪可解释性和公平性嵌入式和RAG分析器LLM跟踪微调Phoenix OSS
Programming255.8K+4.7%
24
METR
METR
METR 是 METR Substack twitter 旗下与 metr.org 对应的产品官网,主要提供:METR is a research nonprofit that evaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose.
General AI253.5K-9.4%
25
Mlflow
Mlflow
MLflow AI Platform 对应 mlflow.org。The largest open source AI engineering platform for agents, LLMs, and ML models. Debug, evaluate, monitor, and optimize your AI applications. Built for teams of all sizes.
General AI230.6K+7.7%
26
Label Studio
Label Studio
适用于所有数据类型的灵活数据注释支持计算机视觉,自然语言处理,语音,声音和视频模型,可自定义的标签和注释模板,通过Webhooks,Python SDK和API与ML / AI管道集成,ML辅助注释与后端集成,连接到云对象存储 (S3和GCP),使用data manager进行高级数据管理,支持多个项目和用户,并获得数据科学家社区的广泛信任
Programming192.8K-6.1%
27
Promptfoo
Promptfoo
Build Secure AI Applications is represented by promptfoo.dev. The AI Security Platform that catches vulnerabilities in development. Trusted by 127 of the Fortune 500 and 300,000+ developers worldwide.
Programming183.4K+16.1%
28
Vectra
Vectra
Vectra AI is represented by vectra.ai. We protect modern networks from modern attacks. Vectra AI sees attackers' every move, connecting the dots across network, identity, and cloud.
General AI172.7K-6.1%
29
Evidentlyai
Evidentlyai
Evidently AI is represented by evidentlyai.com. Ensure your AI is production-ready. Test LLMs and monitor performance across AI applications, RAG systems, and multi-agent workflows. Built on open-source.
General AI166.8K+6.9%
30
PromptQuorum
PromptQuorum
PromptQuorum 是 PromptQuorum 旗下与 promptquorum.com 对应的产品官网,主要提供:Prompt optimization and management across 25+ AI models. Run one prompt, compare outputs, detect hallucinations, and pick the best answer. Free with your API key.
Programming134.1K+235.1%
31
Voxel51
Voxel51
Voxel51 is represented by voxel51.com. Voxel51 empowers multimodal and physical AI builders to unlock visual data insights and maximize model performance. Start building today.
General AI131.7K+14.8%
32
Stanford Agentic Reviewer
Stanford Agentic Reviewer
AI 论文评审 Agent。
Education124.4K+4.9%
33
Getmaxim
Getmaxim
Maxim AI is represented by getmaxim.ai. Evaluate and ship your AI applications with quality, speed, and reliability.
General AI117.2K+7.6%
34
Confident Ai
Confident Ai
Confident AI is represented by confident-ai.com. Confident AI is the AI quality layer for engineers, QA teams, and product leaders. Benchmark, test, and monitor AI systems with research-backed metrics.
General AI116.3K+21.1%
35
Helicone
Helicone
Helicone.ai is represented by helicone.ai. Routing and monitoring for reliable AI apps - the LLMOps platform behind the fastest-growing AI companies.
General AI103.0K-6.2%
36
Meta ARE
Meta ARE
用于 Agent 评估的动态仿真研究平台。
General AI97.3K+33.7%
37
MagicArena
MagicArena
字节跳动推出的视觉 AI 模型对战平台。
General AI84.9K+1.8%
38
Future AGI
Future AGI
Future AGI is represented by futureagi.com. Build self-improving agents. Catch what breaks. Know why. Fix it. Ship smarter every time.
General AI81.6K+16.6%
39
OrcaRouter
OrcaRouter
OrcaRouter 是 OrcaRouter 旗下与 orcarouter.ai 对应的产品官网,主要提供:One OpenAI-compatible AI gateway for production AI — adaptive routing, load balancing, guardrails, agent firewall, observability and governance across 200+ models.
General AI80.8K+22.8%
40
Raindrop
Raindrop
Raindrop is represented by raindrop.ai. Monitor your AI Agent the right way. Get alerted when your agent fails in production, trace exactly what went wrong, and prove your fix worked.
General AI71.1K+3.6%
41
Respan
Respan
Respan 对应 respan.ai。Respan is the LLM engineering platform that unifies observability, evals, prompt optimization, and a unified LLM gateway. Ship reliable AI applications with confidence.
Programming67.9K-1.8%
42
AgentX
AgentX
使用免编码技能构建AI代理聊天机器人使用自定义数据训练AI代理将AI代理部署到网站或消息传递应用程序中
General AI65.1K-13.4%
43
ClearML
ClearML
数据操作数据管理实验管理和可视化模型培训和生命周期管理协作,仪表板和报告模型管理,仓库和版本控制自动化 (CI/CD) 和管道模型服务和监视完全可见的基础架构使用情况自动将环境打包并部署到远程计算机减少计算,硬件和资源成本,以优化性能
Productivity63.3K-14.6%
44
Fiddler
Fiddler
Fiddler AI is represented by fiddler.ai. The Fiddler AI Control Plane provides enterprises with visibility, context, and control across the agentic lifecycle with observability, guardrails, and governance.
Business55.4K+7.3%
45
credo.ai
credo.ai
AI合规和标准AI采用跟踪AI风险管理生成AI保护措施合规监管
General AI53.3K+6.9%
46
LevelAI
LevelAI
联络中心运营自动化
AI Marketing53.0K+22.6%
47
ora
ora
一键可无限创建个性化聊天机器人,无需编码,共享和集成生成的AI,最大的社区生成聊天机器人库
General AI52.8K+2.6%
48
Giskard
Giskard
AI Red Teaming & LLM Security Platform is represented by giskard.ai. Secure AI agents with Giskard’s continuous AI red teaming. Detect vulnerabilities, improve LLM security, and safeguard your AI systems.
General AI50.2K+15.9%
49
Cekura
Cekura
Cekura is represented by cekura.ai. End-to-end testing and observability for Conversational AI. Run pre-production simulations across diverse personas and monitor production conversations to test instruction-following, tool calls, and conversational quality.
Audio48.8K+22.7%
50
Vibe Island
Vibe Island
住在 macOS 灵动岛里的 AI Agent 控制中心
Programming46.7K+55.6%
Browse all 253 AI Observability & Eval products

Related tags

Data month: 2026-07 · Visit figures are third-party traffic estimates aggregated monthly — best for scale and trends. Multiple domains of the same product are merged. See methodology.