Safety·EASYHUB JOURNAL
New Relic announces AI Evaluation to attach LLM-as-a-judge quality and guardrail scores to distributed traces

What changed
New Relic announced AI Evaluation on October 6 as a new AI Observability capability. An asynchronous LLM-as-a-judge service evaluates sampled production telemetry and attaches probabilistic quality scores to deterministic distributed traces. Configurable guardrails can look for hallucinations, prompt injection, jailbreaks, PII leaks, toxicity and bias, while RAG faithfulness, answer relevance and token-cost analysis help teams isolate whether failures originate in prompts, vector retrieval or backend infrastructure. The product also plans a Prompt Playground, versioned golden datasets and controlled experiments. AI Evaluation is scheduled for public preview in November rather than being generally available today.
- Original title
- New Relic Introduces AI Evaluation to Close the Loop Across the Entire Developer-to-Production Lifecycle with Transaction-level Business Impact
- Source
- New Relic · newrelic.com
- Topic
- Safety
- Source month
- 2026-10
This is a concise EasyHub summary of the linked source, not the full report or original reporting. Availability and preview conditions are described in the summary and original.
Summary page published · Editorial information