Users Online

· Guests Online: 14

· Members Online: 0

· Total Members: 285
· Newest Member: Zarfdrilhor

Forum Threads

Newest Threads
No Threads created
Hottest Threads
No Threads created

Latest Articles

AI Evals Test LLM Apps, RAG and Agents Like an Engineer

AI Evals Test LLM Apps, RAG and Agents Like an Engineer
Total Titles: 76
Categories Most Recent Top Rated Popular Courses
DEMOS
04_031 - 4.9 The Cheapest Fix Averaging Both Orders
04 - LLM-as-Judge, and Its Failure Modes - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
04_032 - 4.10 Checkpoint A Judge You Can Trust
04 - LLM-as-Judge, and Its Failure Modes - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_033 - 5.1 Where Eval Cases Actually Come From
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_034 - 5.2 The Golden Set, and How Big It Needs to Be
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_035 - 5.3 Every Incident Becomes a Fixture
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_036 - 5.4 Synthetic Cases Useful, and How They Lie
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_037 - 5.5 Labelling Without Going Insane
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_038 - 5.6 Dataset Drift Your Evals Expire
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_039 - 5.7 Splitting What You Tune On vs What You Trust
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
05_040 - 5.8 Checkpoint A Dataset That Earns Its Keep
05 - Building a Real Eval Dataset - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
06_041 - 6.1 RAG Fails in Two Places, and They Need Diffe...
06 - Evaluating RAG and Agents - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
06_042 - 6.2 Retrieval Metrics That Mean Something
06 - Evaluating RAG and Agents - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
06_043 - 6.3 Faithfulness and Groundedness
06 - Evaluating RAG and Agents - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
06_044 - 6.4 Catching a Hallucination That Cites a Real S...
06 - Evaluating RAG and Agents - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
06_045 - 6.5 Why Agent Evals Are Harder The Trajectory
06 - Evaluating RAG and Agents - AI Evals Test LLM Apps, RAG and Agents Like an Engineer
1 View(s)
No Rating
From: Superadmin
31.05.17
Render time: 6.42 seconds
31,960,768 unique visits