Arize Phoenix and Galileo: OpenTelemetry-Native Observability and Small-Model Judges
Two poles of the 2025-2026 LLM-observability market, and both fix the same two pains. Arize Phoenix: open-source, OpenTelemetry/OTLP-native tracing and evals you can run anywhere and leave without pain. Galileo: a managed AI-reliability platform whose Luna small models score every production call cheaply in real time — replacing 1% sampling with 100% coverage — now folding into Cisco/Splunk enterprise observability.
01.The Problem: Proprietary Gardens and Judges Too Expensive to Trust
You have now met the observability platforms — LangSmith (Topic 289), Braintrust (Topic 290), Weave (Topic 292). Useful, all of them. And two anxieties follow every one.
Anxiety 1: the format is theirs. Your app records every run in the vendor's private trace shape. If you ever want to switch tools — or your compliance team wants your telemetry inside a stack you already own — you are not moving data, you are re-instrumenting your whole app. The instrument becomes the lock-in.
Anxiety 2: judging costs more than the product. The quality check everyone now uses (Topics 290–291) is LLM-as-judge: ask a strong model to score each answer. Strong models are slow and charge per call. Run a frontier judge on every one of your 10,000 daily calls and the monitoring bill can rival the model bill. So teams "solve" it by sampling 1% — which quietly means 99% of your traffic is unwatched.
So the question becomes
Is there a way to record and evaluate LLM apps that is (a) not owned by one vendor, and (b) cheap enough to run on everything, live?
Two products answer those two halves, and they represent the two poles of the 2025–2026 market: Arize Phoenix (open-source, OpenTelemetry-native) and Galileo (managed platform with tiny custom judge models).
OTel as the interchange, two vendor philosophies
OTel as the interchange, two vendor philosophies
Phoenix proves tracing/evals can be open-source and portable via OTLP; Galileo attacks the cost problem with small fine-tuned judges (Luna) at real-time volume — and lands inside enterprise observability via Cisco/Splunk.
Unlock Topic #293: Arize Phoenix and Galileo: OpenTelemetry-Native Observability and Small-Model Judges
You are viewing a preview. The full in-depth technical walkthrough, worked derivations, and code notebooks for this concept, along with self-assessment quizzes, are available with Pro or Lifetime Access.
Failure modes, high-throughput bottlenecks, and real FAANG implementation decisions.
Interactive system topology diagrams, live parameter simulators, and downloadable SVG charts.
Staff-level multiple-choice quiz questions with instant feedback and answer explanations.
Firebase Google authentication automatically syncs your completed topics and quiz scores.
How clear and actionable was this distributed systems breakdown?