Guide
Premium
Intermediate
AI

LLM Evals: How to Test Your AI Application Before You Ship

Learn how to build evals for LLM applications: datasets from real traffic, deterministic checks, LLM-as-judge with binary rubrics, CI integration, and the biases and common mistakes that silently invalidate your metrics. Includes a complete pytest harness ready to use.

44 minutes read
Josué Puig
6 views

Verificando acceso...

Loading comments...

Related Resources

Tutorial

Tutorial: Introduction to LangChain

Learn the basics of LangChain to build AI applications. From installation to your first functional pattern.

Guía
PREMIUM

Context Engineering for AI Agents: Compaction, External Memory, and Subagents

Learn to treat your agent's context window as the finite resource it is: the four levers of context engineering (write, select, compress, isolate), compaction with structured summaries, persistent memory outside the context, just-in-time retrieval, subagents with isolated windows, and tool-result hygiene. With production-ready Python code and the mistakes that silently degrade your agent. New expansion: the context-failure taxonomy (poisoning, distraction, confusion, clash), reasoning budgets with interleaved extended thinking, the real cost of multimodal context, managed memory (Letta, Zep, mem0), effective context length per RULER, and multi-agent handoffs with structured payloads. Latest expansion: the positional anatomy of context (lost in the middle and cache-aware placement), isolating untrusted content by design with Dual-LLM and CaMeL, and generating outputs longer than the window with outlines, rolling summaries, and patch-based revision. August extension: per-section token budgets with a degradation ladder, prompt compression with LLMLingua, semantic retrieval deduplication, self-hosted KV cache (vLLM and SGLang) with cache-aware routing, and memory evaluation with LongMemEval.

Guía

Guide: RAG in Production — Chunking, Embeddings, Hybrid Search, and Reranking

Battle-tested patterns for building production-ready RAG (Retrieval-Augmented Generation) systems in 2026: semantic chunking, embedding selection, hybrid search, and cross-encoder reranking.