AI Evaluation & Assessment

AI evaluation is the discipline of measuring, testing, and validating AI systems. Whether you're building a classifier, fine-tuning an LLM, or deploying a RAG pipeline, rigorous evaluation determines whether your system works and how well it performs in production.

Good evaluation answers critical questions: Does my model generalize? Which metrics matter for my use case? How does it perform on edge cases? Am I introducing bias? Skipping or oversimplifying evaluation is a leading cause of AI system failures in production.

This comprehensive guide covers the full spectrum of evaluation techniques โ€” from fundamental classification metrics (precision, recall, F1) through advanced frameworks like RAGAS for RAG systems and LLM-as-judge patterns for generative models. You'll learn when to use each technique, how to implement them in code, and how to interpret results reliably. By the end, you'll understand how to build robust evaluation pipelines that catch failures before they reach users.

What You'll Learn

Metrics Fundamentals

Master precision, recall, F1-score, accuracy, and confusion matrices. Understand when each metric matters and why accuracy alone is dangerous for imbalanced datasets.

Task-Specific Evaluation

Learn specialized techniques: BLEU/ROUGE for NLP, RAGAS for RAG, ranking metrics for recommendation systems, and loss surfaces for regression.

LLM Evaluation

Evaluate generative models using LLM-as-judge, human preference modeling, and automated scoring frameworks like DeepEval with proper statistical validation.

A/B Testing & Experimentation

Run statistically valid experiments, compute confidence intervals, detect significance, and avoid common statistical mistakes in production deployments.

Why This Matters

Consider these real-world consequences of poor evaluation:

  • A facial recognition system with 99% accuracy overall but 35% error rate on darker skin tones is still deployed, discriminating against users.
  • A classifier trained and tested on balanced data performs terribly on real imbalanced data where the minority class has 1% frequency.
  • An LLM evaluation based solely on perplexity misses hallucination and bias issues that users immediately notice.
  • A model that achieves good performance during development fails when data distribution shifts in production.

Core Philosophy

Evaluation is not a one-time gate โ€” it's continuous. Build evaluation into your development loop from day one. Test on realistic data. Measure what matters for your users, not what's easy to measure. Disaggregate metrics by demographic groups and data domains.

Evaluation in the ML Lifecycle

Evaluation happens at multiple stages:

  • Development: Validate on held-out test set to catch overfitting and verify generalization.
  • Pre-deployment: Run A/B tests with real users or shadow mode to catch distribution shift and fairness issues.
  • Production: Continuously monitor actual user-facing metrics to detect performance degradation and trigger retraining.
  • Post-mortems: When failures occur, analyze evaluation data to understand root causes and prevent recurrence.

Why Evaluation Matters

In production AI systems, a 1% improvement in accuracy might seem minor โ€” until you realize it affects millions of users. Conversely, a model might score well on test metrics but fail catastrophically on real-world data your evaluation didn't capture. This section explains why rigorous evaluation is non-negotiable.

Real Cost of Poor Evaluation

False Confidence

A classifier achieving 95% accuracy on balanced test data might have 5% accuracy on minority classes. Users in those groups experience failure rates that metrics hide.

Distribution Shift

Models trained on 2020 data perform poorly on 2025 data. Evaluation frameworks must catch this drift before production users discover it through degraded service.

Specification Gaming

Optimizing the wrong metric is worse than optimizing nothing. YouTube's watch-time metric led to addictive but low-quality content ranking. Recommenders optimized for engagement promoted misinformation.

Bias & Fairness

Without explicit bias evaluation, systems perpetuate historical discrimination. Facial recognition with 99% accuracy overall but 35% error rate on darker skin tones is still deployed, disproportionately harming minorities.

The Evaluation Mindset

Strong evaluation practices follow these core principles:

  • Measure What Matters: Your primary metric should reflect actual business/user impact, not just statistical convenience. For safety-critical systems, measure failure modes explicitly.
  • Test on Realistic Data: Evaluation data should match production distribution, including edge cases, rare events, and adversarial examples. Don't just test on clean, balanced benchmark data.
  • Use Multiple Metrics: No single metric tells the whole story. Use precision + recall + F1 + AUC + fairness metrics together. Disaggregate by demographic groups, domains, and difficulty levels.
  • Continuous Monitoring: Evaluation doesn't end at deployment. Monitor production performance continuously. Alert when metrics degrade. Retrain when distribution shifts.
  • Adversarial Testing: Actively try to break your model. Test edge cases, domain shifts, and intentional adversarial inputs. Robustness is a feature.
  • Fairness First: Measure bias and fairness explicitly for protected attributes (race, gender, age). Different groups may have vastly different error rates.

The Impact in Numbers

The performance improvement from using better evaluation practices is dramatic:

With Rigorous Eval
96.2%
Standard Eval
84.1%
Minimal Eval
71.3%
No Eval (Prod Fail)
42.0%

Real-world production performance comparison โ€” impact of evaluation rigor on actual system success rates

Why Transformers Won (Applied to Evaluation)

Just as transformers revolutionized NLP through a better architecture, modern evaluation has evolved. Moving from:

  • From: Single accuracy metric To: Disaggregated metrics by group, domain, difficulty
  • From: Evaluation once at end To: Continuous evaluation integrated in development pipeline
  • From: Manual evaluation To: Automated pipelines (RAGAS, DeepEval) with LLM judges
  • From: Binary deployed/not deployed To: Staged rollouts with A/B testing and canary deployments

Key Insight: The model's success isn't determined by accuracy alone โ€” it's determined by evaluation rigor. A mediocre model with excellent evaluation infrastructure often outperforms a great model with poor evaluation. Invest in evaluation first.

Classification Metrics Fundamentals

Classification is the most common ML task: predicting discrete categories (spam/not spam, disease/no disease, sentiment, toxicity level). Evaluating classifiers requires understanding a palette of metrics, each illuminating different aspects of model behavior. Accuracy is just the starting point โ€” real evaluation requires multiple metrics and deep understanding of tradeoffs.

Accuracy: The Seductive Trap

Accuracy is the fraction of correct predictions: (TP + TN) / (TP + TN + FP + FN). It's intuitive and widely used, but often misleading.

Critical Example: In a dataset where 99% of emails are legitimate (imbalanced), a classifier that predicts "not spam" for everything achieves 99% accuracy but is useless. It caught zero spam (recall = 0%), and every spam email reaches users. Users experience 100% failure rate on the task they care about.

This is why accuracy is called "the seductive trap" โ€” it looks good but hides critical failures on the minority class where you likely need the model most.

Accuracy
Accuracy = (TP + TN) / (TP + TN + FP + FN)
TP=True Positives, TN=True Negatives, FP=False Positives, FN=False Negatives
Good for balanced datasets. Dangerous for imbalanced datasets.

Precision & Recall: The Fundamental Tradeoff

Precision answers: "Of the positive predictions I made, how many were correct?"
Recall answers: "Of the actual positives, how many did I catch?"

These metrics are in fundamental tension. You can increase precision by being conservative (predict positive rarely โ€” if you predict positive, you're usually right). But you'll hurt recall because you miss many real positives. You can increase recall by being liberal (predict positive often โ€” you catch most positives). But precision drops because many predictions are false alarms.

Precision Example: A spam classifier that predicts "spam" only for emails matching "Click here now to win $$$". Precision = 99% (almost all flagged emails are spam), but recall = 5% (it misses 95% of spam).

Recall Example: A spam classifier that predicts "spam" for anything containing a link. Recall = 99% (catches nearly all spam), but precision = 20% (80% of flagged emails are legitimate).

Precision & Recall
Precision = TP / (TP + FP)
Recall = TP / (TP + FN) = Sensitivity = True Positive Rate (TPR)
Precision: accuracy of positive predictions. Recall: fraction of actual positives caught.

F1-Score: The Harmonic Mean

F1-score balances precision and recall with a harmonic mean. It penalizes extreme imbalance between the two, forcing both to be reasonably good.

Why harmonic mean? If precision = 99% and recall = 1%, the arithmetic mean is 50%, but the model is nearly useless. The harmonic mean gives ~2%, correctly reflecting that the model fails on recall.

F1-Score
F1 = 2 ร— (Precision ร— Recall) / (Precision + Recall)
If either precision or recall is 0, F1 = 0.
F1 = 1 is perfect. F1 = 0 is total failure.

ROC Curve & AUC-ROC

ROC (Receiver Operating Characteristic) plots TPR (recall) vs FPR (false positive rate) at different classification thresholds. This shows how your classifier trades off between catching positives and avoiding false alarms across all possible decision thresholds.

AUC (Area Under Curve) ranges from 0 to 1. AUC = 0.5 means the classifier is random. AUC = 1.0 means perfect ranking ability.

Key Insight: AUC is threshold-independent. It measures your model's ability to rank positives higher than negatives, regardless of where you set the decision threshold. This makes AUC good for comparing models when you're unsure what threshold to use in production.

FPR in ROC
FPR = FP / (FP + TN) = 1 - Specificity
Specificity = TN / (TN + FP) = true negative rate

When to Use Each Metric

Metric Best For Real Example Warning Accuracy Balanced datasets, all errors equally costly Balanced iris flower classification Hides imbalance problems Precision False positives expensive Loan approval (bad approvals = losses) May miss many positives Recall False negatives expensive Disease detection (missed cases = deaths) May have many false alarms F1 Imbalanced datasets, balanced error cost Fraud detection (both types matter) Doesn't account for actual costs AUC-ROC Threshold-agnostic comparison Comparing two models before deciding threshold Insensitive to class imbalance
Python โ€” Computing Metrics with sklearn
from sklearn.metrics import ( accuracy_score, precision_score, recall_score, f1_score, confusion_matrix, classification_report, roc_auc_score, roc_curve, auc ) import numpy as np y_true = np.array([1, 1, 0, 1, 0, 0, 1, 0, 1, 0]) y_pred = np.array([1, 0, 0, 1, 0, 1, 1, 0, 1, 1]) # Basic metrics acc = accuracy_score(y_true, y_pred) prec = precision_score(y_true, y_pred) rec = recall_score(y_true, y_pred) f1 = f1_score(y_true, y_pred) print(f"Accuracy: {acc:.3f}") print(f"Precision: {prec:.3f}") print(f"Recall: {rec:.3f}") print(f"F1-Score: {f1:.3f}") # Detailed report print("\nClassification Report:") print(classification_report(y_true, y_pred)) # For probabilities / ranking y_scores = np.array([0.9, 0.1, 0.2, 0.8, 0.3, 0.7, 0.95, 0.15, 0.85, 0.25]) auc = roc_auc_score(y_true, y_scores) print(f"\nAUC-ROC: {auc:.3f}") # ROC curve for visualization fpr, tpr, thresholds = roc_curve(y_true, y_scores) roc_auc = auc(fpr, tpr)

Metric Selection Guide

For imbalanced data, NEVER use accuracy alone. Always compute precision, recall, F1, and disaggregate by class. For ranking problems, use AUC-ROC. For threshold-dependent decisions (e.g., loan approval), use precision + recall together or set threshold based on business cost.

Confusion Matrix & Deep Interpretation

The confusion matrix is a 2D grid showing all four outcomes: TP (true positive), FP (false positive), TN (true negative), FN (false negative). Visualizing it reveals patterns that aggregate metrics hide. Understanding the confusion matrix is the key to understanding your classifier's behavior.

Binary Confusion Matrix Anatomy

Predicted +
Predicted โˆ’
Actual +
TP
True Positive
Correct detection
FN
False Negative
Missed positive
Recall = TP/(TP+FN)
Actual โˆ’
FP
False Positive
False alarm
TN
True Negative
Correct rejection
Specificity = TN/(TN+FP)

Interpreting Each Cell

  • TP (Top-Left): We predicted positive and were right. These are successful positive predictions. More TP is better.
  • FN (Top-Right): We predicted negative but were wrong. We missed actual positives. This is the cost of low recall. High FN means missing real cases.
  • FP (Bottom-Left): We predicted positive but were wrong. False alarms. This is the cost of low precision. High FP means wasting resources on false leads.
  • TN (Bottom-Right): We predicted negative and were right. Successful negative predictions. More TN is better.

Visual Pattern Recognition

The confusion matrix shape tells you about your model's behavior:

  • Diagonal strong (TP & TN large): Good classifier, balanced predictions.
  • Top-right large (FN large): Model misses positives. Recall is low. Too conservative.
  • Bottom-left large (FP large): Model has false alarms. Precision is low. Too liberal.
  • Bottom-right huge (TN dominates): Imbalanced data with majority class prediction bias.
Python โ€” Confusion Matrix Visualization
from sklearn.metrics import confusion_matrix import matplotlib.pyplot as plt import seaborn as sns import numpy as np # Example predictions y_true = [1, 1, 0, 1, 0, 0, 1, 0, 1, 0, 1, 1, 0, 1, 0] y_pred = [1, 0, 0, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 0, 0] # Compute confusion matrix cm = confusion_matrix(y_true, y_pred) # Visualize fig, ax = plt.subplots(figsize=(8, 6)) sns.heatmap(cm, annot=True, fmt='d', cmap='Blues', xticklabels=['Negative', 'Positive'], yticklabels=['Negative', 'Positive'], ax=ax, cbar_kws={'label': 'Count'}) ax.set_ylabel('Actual Label') ax.set_xlabel('Predicted Label') ax.set_title('Confusion Matrix') # Extract and print metrics from confusion matrix tn, fp, fn, tp = cm.ravel() print(f"TP: {tp}, FP: {fp}") print(f"FN: {fn}, TN: {tn}") # Compute derived metrics sensitivity = tp / (tp + fn) if (tp + fn) > 0 else 0 # Recall specificity = tn / (tn + fp) if (tn + fp) > 0 else 0 precision = tp / (tp + fp) if (tp + fp) > 0 else 0 f1 = 2 * (precision * sensitivity) / (precision + sensitivity) if (precision + sensitivity) > 0 else 0 print(f"\nMetrics from Confusion Matrix:") print(f"Sensitivity (Recall): {sensitivity:.3f}") print(f"Specificity: {specificity:.3f}") print(f"Precision: {precision:.3f}") print(f"F1-Score: {f1:.3f}") plt.tight_layout() plt.savefig('confusion_matrix.png', dpi=150) plt.show()

Multi-Class Confusion Matrix

For multi-class (3+ classes), the matrix grows to Nร—N. Visualize it as a heatmap (easier to read than a table). Compute metrics per-class and macro/weighted averages. Look for which class pairs the model confuses most.

Regression Metrics

Regression predicts continuous values (house prices, temperature, stock returns, user engagement scores). Unlike classification where predictions are right/wrong, regression allows varying degrees of closeness. Different regression metrics penalize errors differently โ€” choosing the right metric is crucial.

Mean Absolute Error (MAE)

MAE is the average absolute difference between predicted and actual values. It's intuitive and in the same units as your target. Treat each error equally regardless of magnitude.

Interpretation: If MAE = $50,000 on house prices, predictions are off by $50k on average.

MAE
MAE = (1/n) ร— ฮฃ|y_i - ลท_i|
Units same as target. Treats all errors equally. Robust to outliers.

Mean Squared Error (MSE) & RMSE

MSE squares errors before averaging, heavily penalizing large mistakes. If one prediction is off by 100 and another by 1, MSE cares much more about the first. RMSE is the square root of MSE, converting back to original units.

When to use MSE/RMSE: When large errors are much worse than small errors (compound cost). Stock price errors of $100 are much worse than $1 errors.

MSE & RMSE
MSE = (1/n) ร— ฮฃ(y_i - ลท_i)ยฒ
RMSE = โˆšMSE
Differentiable (good for optimization). Sensitive to outliers. Units same as target.

Rยฒ Score (Coefficient of Determination)

Rยฒ measures the fraction of variance in y explained by the model. Rยฒ = 1 means perfect prediction. Rยฒ = 0 means the model is no better than predicting the mean y_avg for everything. Rยฒ < 0 means the model is worse than the mean.

Interpretation: Rยฒ = 0.85 means "our model explains 85% of the variation in the data".

Rยฒ
Rยฒ = 1 - (SS_res / SS_tot)
SS_res = ฮฃ(y_i - ลท_i)ยฒ (residual sum of squares)
SS_tot = ฮฃ(y_i - ศณ)ยฒ (total sum of squares)
Normalized [-โˆž, 1]. Interpretable percentage.

Comparing Regression Metrics

Metric Range When to Use Key Property MAE [0, โˆž] Outliers present, errors equally important Robust, interpretable MSE [0, โˆž] For optimization (differentiable) Penalizes large errors quadratically RMSE [0, โˆž] Large errors very bad, same units as target Interpretable, sensitive to outliers Rยฒ [-โˆž, 1] Communicating model quality to non-technical stakeholders Normalized, interpretable percentage MAPE [0, โˆž]% Percentage errors (sales forecasting) Scale-independent, interpretable
Python โ€” Regression Metrics
from sklearn.metrics import ( mean_absolute_error, mean_squared_error, r2_score, mean_absolute_percentage_error ) import numpy as np y_true = np.array([3.0, 2.5, 2.0, 3.5, 2.8, 4.1, 2.9]) y_pred = np.array([3.1, 2.4, 2.2, 3.3, 3.0, 3.9, 2.7]) # Compute metrics mae = mean_absolute_error(y_true, y_pred) mse = mean_squared_error(y_true, y_pred) rmse = np.sqrt(mse) r2 = r2_score(y_true, y_pred) mape = mean_absolute_percentage_error(y_true, y_pred) print("=== Regression Metrics ===") print(f"MAE: {mae:.4f} (avg absolute error)") print(f"MSE: {mse:.4f} (squared error)") print(f"RMSE: {rmse:.4f} (root mean squared error)") print(f"Rยฒ: {r2:.4f} ({r2*100:.1f}% variance explained)") print(f"MAPE: {mape:.2%} (percentage error)") # Residuals analysis residuals = y_true - y_pred print(f"\nResidual Stats:") print(f"Mean residual: {residuals.mean():.4f} (should be ~0)") print(f"Std residual: {residuals.std():.4f}") print(f"Min/Max: {residuals.min():.4f} / {residuals.max():.4f}")

Residual Analysis

Don't just compute aggregate metrics. Plot residuals (actual - predicted) vs predicted values. Should be randomly scattered around zero with no patterns. Patterns indicate model bias or heteroscedasticity.

NLP-Specific Metrics: BLEU & ROUGE

For language tasks (machine translation, summarization, question-answering), traditional metrics (accuracy, MSE) don't apply. Instead, we use specialized metrics that compare predicted text to reference text. These metrics measure n-gram overlap and sequence similarity.

BLEU (Bilingual Evaluation Understudy)

BLEU measures n-gram precision: what fraction of n-grams in the predicted text also appear in the reference text? Originally designed for machine translation evaluation. Emphasis on precision: are the words the model outputs correct?

BLEU scores range from 0 to 1 (or 0-100 when scaled). Common thresholds: BLEU > 0.3 is usually acceptable, > 0.4 is good, > 0.5 is very good.

Example: Reference: "the quick brown fox jumps" | Prediction: "the fast brown fox runs" | BLEU captures that "the", "brown", "fox" are shared (3 unigrams match), but "quick" โ‰  "fast" and "jumps" โ‰  "runs".

BLEU Concept
BLEU = exp( ฮฃ w_n log(p_n) )
where p_n = min(#n-grams in pred โˆฉ reference) / (#n-grams in pred)
weights w_n typically (0.25, 0.25, 0.25, 0.25) for 1-4 grams

ROUGE (Recall-Oriented Understudy for Gisting Evaluation)

ROUGE focuses on recall: what fraction of reference n-grams appear in predicted text? Particularly good for summarization evaluation because it measures content preservation.

Common variants:

  • ROUGE-1 (Unigrams): Single word overlap. Captures basic content.
  • ROUGE-2 (Bigrams): Two-word phrase overlap. Captures more context.
  • ROUGE-L (Longest Common Subsequence): Longest matching sequence. Captures word order.
Python โ€” BLEU & ROUGE Computation
from nltk.translate.bleu_score import sentence_bleu, SmoothingFunction from rouge_score import rouge_scorer import nltk nltk.download('punkt') # Example: Machine translation evaluation reference = "the cat is on the mat" hypothesis = "the cat is on mat" # BLEU score (nltk) ref_tokens = reference.split() hyp_tokens = hypothesis.split() bleu = sentence_bleu([ref_tokens], hyp_tokens, weights=(0.25, 0.25, 0.25, 0.25), smoothing_function=SmoothingFunction().method1) print(f"BLEU: {bleu:.4f}") # ROUGE score (rouge_score library) scorer = rouge_scorer.RougeScorer(['rouge1', 'rouge2', 'rougeL'], use_stemmer=True) scores = scorer.score(reference, hypothesis) print("\nROUGE Scores:") for metric, score in scores.items(): print(f"{metric}:") print(f" Precision: {score.precision:.3f}") print(f" Recall: {score.recall:.3f}") print(f" F-measure: {score.fmeasure:.3f}")

BLEU vs ROUGE: When to Use Each

Use BLEU For

Machine translation, paraphrase detection, any task where word-for-word accuracy matters. Emphasizes precision: are outputs grammatically/semantically correct?

Use ROUGE For

Summarization, extraction tasks, where content preservation matters more than exact wording. Emphasizes recall: did you capture the important content?

Use LLM-Judge For

Any task where semantic quality matters. BLEU/ROUGE often miss meaning. A gold-standard answer might score 0 BLEU but have perfect meaning.

Limitations of BLEU/ROUGE

These metrics are based on n-gram overlap, which is crude. Two paraphrases with identical meaning might have 0 BLEU overlap if they use different words. For modern systems, consider pairing with LLM-as-judge for semantic evaluation. Many papers show BLEU/ROUGE correlate poorly with human judgment.

RAG Evaluation with RAGAS

Retrieval-Augmented Generation (RAG) systems combine two stages: retrieve relevant documents, then generate answers grounded in those documents. Evaluating RAG requires assessing both retrieval quality AND generation quality, plus whether they work together correctly.

The RAG Pipeline

RAG = Retrieval (finding relevant documents) + Augmentation (feeding them to generator) + Generation (creating answers). A RAG system can fail at any stage: retrieve irrelevant docs, pass relevant docs but generator ignores them, or generate hallucinated content.

RAGAS Framework

RAGAS (Retrieval-Augmented Generation Assessment) provides automated metrics without requiring human annotation. It uses LLMs to evaluate dimensions of RAG quality.

Core RAGAS metrics:

  • Faithfulness (0-1): Is the generated answer grounded in the retrieved context? Penalizes hallucinations. Uses LLM to check if answer claims appear in context.
  • Answer Relevance (0-1): Does the answer address the question well? Semantic alignment. Uses embeddings/LLM to verify relevance.
  • Context Relevance (0-1): Are retrieved documents actually relevant to the question? Assesses retriever quality directly.
  • Context Recall (0-1): Do retrieved documents contain enough information to answer the question? Can the question be answered from the context?
  • Context Precision (0-1): Is there minimal irrelevant information in retrieved context? Measures signal-to-noise ratio.
Metric What It Measures Target Value Failure Mode Faithfulness Answer grounded in context > 0.8 Hallucinations (facts not in docs) Answer Relevance Answer addresses question > 0.8 Off-topic or incomplete answers Context Relevance Retrieved docs relevant > 0.7 Retriever returns irrelevant docs Context Recall Sufficient info in docs > 0.7 Information need not covered Context Precision Minimal noise in context > 0.7 Too many irrelevant passages
Python โ€” RAGAS Evaluation
from ragas import evaluate from ragas.metrics import ( faithfulness, answer_relevance, context_relevance, context_recall, context_precision ) from datasets import Dataset # Prepare your RAG eval data # Each sample needs: question, answer, context (list of doc chunks), ground_truth data = { "question": [ "What is photosynthesis?", "How does DNA replication work?" ], "answer": [ "Photosynthesis is the process where plants convert sunlight into chemical energy...", "DNA replication occurs through semiconservative copying via DNA polymerase..." ], "context": [ [ "Plants use sunlight to create energy through photosynthesis.", "Chlorophyll absorbs light energy and drives electron transport chains.", "The light reactions and Calvin cycle together convert CO2 to glucose." ], [ "DNA is a double helix molecule composed of nucleotides.", "Helicases unwind the DNA strand, and DNA polymerase synthesizes new strands.", "The process is semiconservative: each new DNA molecule has one old and one new strand." ] ], "ground_truth": [ "Photosynthesis converts light energy into chemical energy stored in glucose", "DNA replication creates two identical DNA molecules from one original" ] } dataset = Dataset.from_dict(data) # Run evaluation results = evaluate( dataset, metrics=[ faithfulness, answer_relevance, context_relevance, context_recall, context_precision ] ) # Analyze results print("=== RAGAS Results ===") print(f"Faithfulness: {results['faithfulness'].mean():.3f}") print(f"Answer Relevance: {results['answer_relevance'].mean():.3f}") print(f"Context Relevance: {results['context_relevance'].mean():.3f}") print(f"Context Recall: {results['context_recall'].mean():.3f}") print(f"Context Precision: {results['context_precision'].mean():.3f}") # Per-sample analysis for i, row in results.to_pandas().iterrows(): print(f"\nSample {i}:") if row['faithfulness'] < 0.8: print(f" WARNING: Low faithfulness ({row['faithfulness']:.2f}) - possible hallucination") if row['context_recall'] < 0.7: print(f" WARNING: Low context recall ({row['context_recall']:.2f}) - missing info")

RAG Evaluation Best Practice

Don't just measure answer quality in isolation. The RAGAS metrics tell a complete story: retrieval quality (context relevance/recall/precision), generation quality (answer relevance), and grounding (faithfulness). A 0.9 answer relevance with 0.3 faithfulness is worse than a 0.7 relevance with 0.95 faithfulness.

LLM-as-Judge Pattern

For generative models (large language models, summarizers, dialogue systems), traditional metrics fall short. BLEU/ROUGE are n-gram based and miss semantic quality. Human evaluation is expensive. The emerging solution: use a powerful LLM as a judge to evaluate outputs along custom rubrics.

How LLM-as-Judge Works

You provide: 1) Evaluation criteria, 2) The question/prompt, 3) The generated output, 4) Optionally a reference answer. The LLM returns: a score and explanation. Essentially you're asking "Does this output satisfy this rubric?"

Advantages & Limitations

Advantages

Semantic evaluation captures nuance traditional metrics miss. Flexible custom criteria. Cheap at scale (LLM call per eval). Studies show LLM judgments correlate better with human consensus than BLEU.

Limitations

LLMs have systematic biases. Can hallucinate explanations. Reproducibility varies. Requires careful prompt engineering. Slower than automatic metrics. Cost accumulates on large eval sets.

Python โ€” LLM-as-Judge with DeepEval
from deepeval.metrics import Faithfulness, Relevance, Coherence from deepeval import evaluate # DeepEval handles prompt engineering for you faithfulness = Faithfulness() relevance = Relevance() coherence = Coherence() # Prepare test cases test_cases = [ { "question": "What is machine learning?", "output": "Machine learning is a subset of artificial intelligence that enables systems to learn...", "retrieval_context": ["ML algorithms learn patterns from data without explicit programming..."] }, { "question": "Explain neural networks", "output": "Neural networks are computing systems inspired by biological neurons...", "retrieval_context": ["An artificial neural network consists of interconnected nodes..."] } ] # Evaluate results = evaluate(test_cases, metrics=[faithfulness, relevance, coherence]) # Results contain scores and LLM explanations for i, result in enumerate(results): print(f"\nTest Case {i}:") print(f" Faithfulness: {result['faithfulness'].score}/1.0") print(f" Reason: {result['faithfulness'].reason}") print(f" Relevance: {result['relevance'].score}/1.0")

Prompting Best Practices for LLM Judges

  • Be Specific with Rubrics: Vague criteria produce inconsistent scores. Define "good" precisely with examples.
  • Use Structured Scales: Define what 5/5, 3/3, or 1/10 means. Provide anchor points (1=poor, 3=acceptable, 5=excellent).
  • Provide Context: Include the question, generated answer, and reference(s). More context reduces hallucination.
  • Request Reasoning: Ask the judge LLM to explain its score. Reasoning increases reliability and helps debug failures.
  • Use Consistent LLM: Judge with the same model/temperature. Different LLMs judge differently.
  • Validate with Humans: For critical domains, compare LLM judgments against human consensus. Calibrate if they diverge.
  • Test for Bias: Check if judge correlates with protected attributes (e.g., does it systematically score certain demographic outputs lower?).

LLM Judge Prompt Template

You are an expert evaluator. Rate the following answer on a scale of 1-5: 1 = Completely wrong or hallucinated 3 = Partially correct, some hallucinations or irrelevance 5 = Accurate, relevant, grounded in the context Question: {question} Context: {context} Generated Answer: {answer} Provide: 1) Your score, 2) Specific evidence, 3) Explanation.

A/B Testing & Statistical Significance

A/B testing compares two models (or strategies) by deploying them to different user groups and measuring which performs better on real usage metrics. Proper A/B testing requires statistical rigor to avoid false positives (claiming improvement that doesn't exist) and false negatives (missing real improvements).

Key Concepts

  • Null Hypothesis (Hโ‚€): There is NO difference between A and B (any observed difference is noise).
  • Alternative Hypothesis (Hโ‚): There IS a difference between A and B (improvement is real).
  • P-value: Probability of observing this data if Hโ‚€ is true. Lower p-value = stronger evidence against Hโ‚€. If p < 0.05, we reject Hโ‚€.
  • Significance Level (ฮฑ): Usually 0.05 (5% false positive rate). If p < ฮฑ, we declare the result statistically significant.
  • Type I Error (False Positive): Declaring improvement when none exists (probability = ฮฑ).
  • Type II Error (False Negative): Missing real improvement (probability = ฮฒ). Power = 1 - ฮฒ.
  • Power: Probability of detecting a real improvement. Target power โ‰ฅ 0.8 (80%).
Effect Size (Cohen's d)
Cohen's d = (mean_A - mean_B) / pooled_std
0.2 = small effect, 0.5 = medium, 0.8 = large
Always report effect size alongside p-values
Python โ€” A/B Testing with Statistical Significance
from scipy import stats import numpy as np from math import sqrt # Simulate: 2000 users in group A (model A), 2000 in group B (model B) # Metric: click-through rate (CTR) np.random.seed(42) group_a = np.random.binomial(1, p=0.12, size=2000) # 12% CTR group_b = np.random.binomial(1, p=0.135, size=2000) # 13.5% CTR (2.5% improvement) ctr_a = group_a.mean() ctr_b = group_b.mean() improvement = (ctr_b - ctr_a) / ctr_a * 100 print(f"Group A: CTR = {ctr_a:.3%} ({group_a.sum()} conversions / {len(group_a)} users)") print(f"Group B: CTR = {ctr_b:.3%} ({group_b.sum()} conversions / {len(group_b)} users)") print(f"Improvement: {improvement:.1f}%") # Two-sample t-test t_stat, p_value = stats.ttest_ind(group_a, group_b) print(f"\nStatistical Test Results:") print(f"T-statistic: {t_stat:.4f}") print(f"P-value: {p_value:.6f}") # Determine significance alpha = 0.05 if p_value < alpha: print(f"โœ“ SIGNIFICANT (p = {p_value:.6f} < {alpha})") print(f" Reject Hโ‚€: Difference is real, not due to chance.") else: print(f"โœ— NOT SIGNIFICANT (p = {p_value:.6f} >= {alpha})") print(f" Fail to reject Hโ‚€: Cannot rule out that difference is random.") # Effect size (Cohen's d) pooled_std = sqrt(((len(group_a)-1)*group_a.std()**2 + (len(group_b)-1)*group_b.std()**2) / (len(group_a) + len(group_b) - 2)) cohens_d = (ctr_b - ctr_a) / pooled_std print(f"\nEffect Size (Cohen's d): {cohens_d:.4f}") if abs(cohens_d) < 0.2: effect = "small" elif abs(cohens_d) < 0.5: effect = "small-to-medium" elif abs(cohens_d) < 0.8: effect = "medium-to-large" else: effect = "large" print(f"Interpretation: {effect} effect") # Confidence interval for difference se_diff = sqrt(ctr_a*(1-ctr_a)/len(group_a) + ctr_b*(1-ctr_b)/len(group_b)) ci_lower = (ctr_b - ctr_a) - 1.96 * se_diff ci_upper = (ctr_b - ctr_a) + 1.96 * se_diff print(f"\n95% CI for difference: [{ci_lower:.4f}, {ci_upper:.4f}]") print(f"Real improvement is likely between {ci_lower*100:.2f}% and {ci_upper*100:.2f}%")

Sample Size Calculation (Critical!)

Running an A/B test on 100 users each is likely to be inconclusive (low power). Use power analysis BEFORE running the test to determine required sample size.

Python โ€” Calculating Required Sample Size
from scipy.stats import norm import numpy as np def sample_size_for_proportion(p1, p2, alpha=0.05, power=0.8): """ Calculate sample size per group for binary outcome (proportion) tests. Args: p1: Baseline proportion (group A) p2: Target proportion (group B) alpha: Significance level (false positive rate) power: Power = 1 - false negative rate """ z_alpha = norm.ppf(1 - alpha/2) # Two-tailed z_beta = norm.ppf(power) # One-tailed p_bar = (p1 + p2) / 2 n = 2 * (z_alpha + z_beta)**2 * p_bar * (1 - p_bar) / (p1 - p2)**2 return int(np.ceil(n)) # Example: Current CTR = 10%, want to detect 12% (2% absolute, 20% relative lift) baseline_ctr = 0.10 target_ctr = 0.12 n = sample_size_for_proportion(baseline_ctr, target_ctr, alpha=0.05, power=0.8) total_users = 2 * n print(f"Baseline CTR: {baseline_ctr:.1%}") print(f"Target CTR: {target_ctr:.1%}") print(f"Required users per group: {n:,}") print(f"Total users needed: {total_users:,}") print(f"At 10k users/day, runtime: {total_users/10000:.1f} days") # Sensitivity: what effect sizes can we detect? print("\nDetectable effect sizes with different sample sizes:") for n_test in [500, 1000, 2000, 5000]: # Solve for detectable difference min_diff = norm.ppf(0.975) * np.sqrt(0.1*0.9/n_test + 0.1*0.9/n_test) print(f" n={n_test}: Can detect {min_diff*100:.2f}% difference")

A/B Testing Pitfalls to Avoid

1) Peeking at results early (invalid p-values). 2) Running multiple simultaneous tests without correction. 3) Stopping test when you see 'significance' (p-hacking). 4) Assuming correlation = causation. 5) Ignoring novelty/primacy effects (users may prefer change simply because it's new). Always pre-register your metrics, run at planned sample size, and validate with follow-up tests.

Building Custom Evaluation Pipelines

Production ML systems require custom evaluation pipelines tailored to specific business requirements. Building a robust, automated evaluation system that runs on every model iteration is essential for reliable deployment and continuous improvement.

Pipeline Architecture

A typical evaluation pipeline: 1) Load test data, 2) Run model on test data, 3) Compute metrics (multiple, disaggregated), 4) Compare to baseline, 5) Check for regressions, 6) Generate report, 7) Make decision (deploy/reject/investigate).

Python โ€” Building an Evaluation Pipeline
import json from pathlib import Path from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score from datetime import datetime class EvaluationPipeline: """Automated evaluation pipeline for ML models.""" def __init__(self, model, test_data, baseline_metrics=None): self.model = model self.test_data = test_data self.baseline = baseline_metrics or {} self.results = {} def run_evaluation(self, y_true, y_pred, sample_weights=None): """Compute comprehensive metrics.""" metrics = { 'timestamp': datetime.now().isoformat(), 'accuracy': accuracy_score(y_true, y_pred), 'precision': precision_score(y_true, y_pred, average='weighted', zero_division=0), 'recall': recall_score(y_true, y_pred, average='weighted', zero_division=0), 'f1': f1_score(y_true, y_pred, average='weighted', zero_division=0), } # Disaggregate by group if 'group' in self.test_data.columns: metrics['by_group'] = {} for group in self.test_data['group'].unique(): mask = self.test_data['group'] == group group_metrics = { 'accuracy': accuracy_score(y_true[mask], y_pred[mask]), 'f1': f1_score(y_true[mask], y_pred[mask], average='weighted'), 'samples': mask.sum() } metrics['by_group'][str(group)] = group_metrics return metrics def check_regressions(self, metrics): """Alert if metrics degrade vs baseline.""" alerts = [] for metric, value in metrics.items(): if metric in self.baseline and isinstance(value, (int, float)): if value < self.baseline[metric] * 0.95: # 5% drop tolerance alerts.append(f"WARNING: {metric} degraded from {self.baseline[metric]:.3f} to {value:.3f}") return alerts def save_results(self, filepath): """Save evaluation results to JSON.""" with open(filepath, 'w') as f: json.dump(self.results, f, indent=2) def generate_report(self): """Generate human-readable report.""" report = "=== EVALUATION REPORT ===\n" for metric, value in self.results.items(): if isinstance(value, (int, float)): report += f"{metric}: {value:.4f}\n" return report # Usage from sklearn.datasets import load_iris from sklearn.ensemble import RandomForestClassifier from sklearn.model_selection import train_test_split import pandas as pd X, y = load_iris(return_X_y=True) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2) model = RandomForestClassifier() model.fit(X_train, y_train) y_pred = model.predict(X_test) test_df = pd.DataFrame(X_test) test_df['group'] = y_test % 2 # Dummy grouping pipeline = EvaluationPipeline( model, test_df, baseline_metrics={'accuracy': 0.95, 'f1': 0.94} ) metrics = pipeline.run_evaluation(y_test, y_pred) pipeline.results = metrics alerts = pipeline.check_regressions(metrics) if alerts: for alert in alerts: print(alert) print(pipeline.generate_report())

Continuous Evaluation in CI/CD

Integrate evaluation into your ML pipeline so every model change is automatically tested. This catches regressions early.

Python โ€” GitHub Actions Evaluation Workflow
name: Evaluate Model Changes on: [pull_request, push] jobs: evaluate: runs-on: ubuntu-latest steps: - uses: actions/checkout@v3 - uses: actions/setup-python@v4 with: python-version: "3.10" - name: Install dependencies run: pip install -r requirements.txt - name: Run evaluation run: python evaluate.py --output eval_results.json - name: Check metrics meet thresholds run: | python -c " import json with open('eval_results.json') as f: results = json.load(f) assert results['accuracy'] >= 0.85, f"Accuracy too low: {results['accuracy']}" assert results['f1'] >= 0.82, f"F1 too low: {results['f1']}" print('โœ“ All metrics pass thresholds') " - name: Upload results if: always() uses: actions/upload-artifact@v3 with: name: eval-results path: eval_results.json

Bias, Fairness & Responsible Evaluation

Models can achieve high aggregate accuracy while failing systematically on minority groups. A recruitment AI might have 95% accuracy overall but 5% accuracy on applicants from certain backgrounds. Responsible evaluation requires measuring performance across demographic groups and checking for unfair disparities.

Key Fairness Concepts

  • Demographic Parity: Model predictions should have equal rates across protected groups. May not be appropriate: suppose predicting "admits to program" โ€” you might want meritocracy, not equal admission rates across groups.
  • Equalized Odds: True positive rate AND false positive rate should be equal across groups. Stronger: equal error rates for all. More appropriate for most domains.
  • Calibration: Predicted probabilities should be accurate for all groups. A 0.7 confidence prediction should have ~70% true positive rate in all groups, not just overall.
  • Individual Fairness: Similar individuals should receive similar treatment. If two applicants are identical except for protected attribute, they should have same prediction.
Python โ€” Fairness Audit
import pandas as pd from sklearn.metrics import confusion_matrix, precision_score, recall_score # Suppose we have predictions and demographic groups df = pd.DataFrame({ 'y_true': [1, 1, 0, 1, 0, 0, 1, 0, 1, 0, 1, 1, 0, 1, 0, 1, 0, 0, 1, 0], 'y_pred': [1, 0, 0, 1, 0, 1, 1, 0, 1, 1, 1, 1, 0, 0, 0, 1, 1, 0, 1, 0], 'group': ['A']*10 + ['B']*10 # e.g., age groups, ethnicities }) print("=== FAIRNESS AUDIT ===\n") for group in df['group'].unique(): subset = df[df['group'] == group] # Compute confusion matrix tn, fp, fn, tp = confusion_matrix( subset['y_true'], subset['y_pred'] ).ravel() # Compute fairness metrics tpr = tp / (tp + fn) if (tp + fn) > 0 else 0 # True positive rate (recall) fpr = fp / (fp + tn) if (fp + tn) > 0 else 0 # False positive rate precision = tp / (tp + fp) if (tp + fp) > 0 else 0 print(f"Group {group} (n={len(subset)}):") print(f" TPR (Sensitivity): {tpr:.3f}") print(f" FPR (1-Specificity): {fpr:.3f}") print(f" Precision: {precision:.3f}") print(f" Samples: {len(subset)}") print() # Check for disparate impact (hiring discrimination proxy) print("\nDisparate Impact Analysis:") selection_a = (df[df['group']=='A']['y_pred']==1).mean() selection_b = (df[df['group']=='B']['y_pred']==1).mean() di_ratio = selection_b / selection_a if selection_a > 0 else 0 print(f"Group A selection rate: {selection_a:.1%}") print(f"Group B selection rate: {selection_b:.1%}") print(f"Disparate Impact Ratio (B/A): {di_ratio:.3f}") if di_ratio < 0.8: print("โš  WARNING: Ratio < 0.8 indicates potential discrimination (EEO 4/5 rule)") elif di_ratio > 1.25: print("โš  WARNING: Ratio > 1.25 indicates potential reverse discrimination") else: print("โœ“ Ratio within acceptable range [0.8, 1.25]")

Fairness Tradeoffs

Different fairness definitions can be mathematically incompatible. You cannot simultaneously satisfy demographic parity AND equalized odds in imbalanced datasets with unequal base rates. Choose your fairness criteria based on context and stakeholder input, then evaluate explicitly. Document your choices.

Evaluation Best Practices

1. Data Splitting Strategy

  • Stratified Split: For imbalanced datasets, stratify on the target to keep class distributions equal across train/val/test.
  • Time-Based Split: For time-series data, always test on future data (never test on past). Training on 2024, testing on 2025.
  • Domain-Based Split: For multi-domain tasks, evaluate each domain separately. A model may work on Twitter but fail on Reddit.
  • Hold-Out Test Set: Never, ever use test data for any tuning. Use completely separate held-out set for final evaluation, opened only once.
  • Nested Cross-Validation: Inner loop tunes hyperparameters, outer loop reports final performance. This prevents leakage.

2. Evaluation Data Requirements

  • Representative Distribution: Test data should match production distribution. If production is 1% positive and test is 50/50, results are misleading.
  • Sufficient Size: Use power analysis to ensure test sets are large enough. 50 samples per class is often insufficient for reliable estimates.
  • Multiple Datasets: Evaluate on multiple independent datasets. A model may overfit to one benchmark.
  • Edge Cases: Deliberately include edge cases, rare events, boundary conditions. These are where most failures occur.
  • Adversarial Examples: Include examples designed to break the model. Robustness is a feature, not a bug.

3. Metric Selection

  • Primary Metric: Choose ONE primary metric aligned with business objectives (e.g., revenue impact, user safety, fairness). This is your north star.
  • Secondary Metrics: Track 2-3 secondary metrics to catch regressions in other dimensions. For example: primary=F1, secondary=[precision, recall, fairness].
  • Disaggregated Metrics: Always disaggregate by demographic groups, data domains, and difficulty levels. Aggregate metrics hide failures.
  • Human Baselines: Measure human performance when available. Your model should beat humans on the task you're solving.

4. Continuous Production Monitoring

  • Production Metrics: Monitor actual user-facing metrics continuously (not just offline metrics). CTR, conversion rate, retention.
  • Distribution Monitoring: Alert when input distribution changes (data drift). If your training data was 50% male, 50% female, but current production is 80% male, your model will fail.
  • Performance Regression: Alert when metrics decline below acceptable thresholds. Have clear escalation procedures.
  • Automated Retraining: Automatically retrain when performance drops. Only deploy if automated eval passes.

Common Evaluation Pitfalls & How to Avoid Them

Pitfall: Accuracy Only

Accuracy is meaningless on imbalanced data. 99% accurate classifier on 99%-negative data is useless. Use precision+recall+F1+disaggregation.

Pitfall: Test Leakage

If you train on data similar to test set, results are artificially high. Strictly separate train/val/test. Don't tune on test data.

Pitfall: Missing Baselines

Comparing to nothing is meaningless. Always baseline: 1) random, 2) simple heuristic, 3) human, 4) prior work.

Pitfall: Static Evaluation

Evaluate once, deploy, never test again โ€” recipe for disaster. Continuously monitor and retrain on production data.

Pitfall: Ignoring Edge Cases

Benchmarks hide failures on rare cases. Deliberately test edge cases, adversarial examples, out-of-distribution inputs.

Pitfall: Single Metric Opt

Optimizing one metric often degrades others. YouTube's watch-time metric promoted addictive trash. Measure multiple objectives.

Pitfall: P-Hacking & Multiple Comparisons

Run enough statistical tests and you'll find "significance" by chance. Always pre-register metrics before looking at data. If testing 10 metrics at ฮฑ=0.05, expect ~0.5 false positives by chance alone.

Python โ€” Correcting for Multiple Comparisons
from scipy.stats import ttest_ind from statsmodels.stats.multitest import multipletests import numpy as np # Testing 10 different metrics (dangerous!) np.random.seed(42) metrics_p_values = [] for metric_id in range(10): group_a = np.random.randn(100) group_b = np.random.randn(100) + 0.05 # Tiny effect t_stat, p_value = ttest_ind(group_a, group_b) metrics_p_values.append(p_value) print("Unadjusted significance (p < 0.05):") unadjusted_sig = sum(p < 0.05 for p in metrics_p_values) print(f" {unadjusted_sig} metrics significant (expect ~0.5 false positives)") # Apply Bonferroni correction: divide threshold by number of tests bonferroni_threshold = 0.05 / len(metrics_p_values) bonferroni_sig = sum(p < bonferroni_threshold for p in metrics_p_values) print(f"\nBonferroni correction (threshold = {bonferroni_threshold:.4f}):") print(f" {bonferroni_sig} metrics significant (conservative)") # Better: Benjamini-Hochberg FDR control (less conservative, controls false discovery rate) rejected, adjusted_p, sidak, bonf = multipletests(metrics_p_values, method='fdr_bh') fdr_sig = sum(rejected) print(f"\nBenjamini-Hochberg FDR (controls false discovery rate):") print(f" {fdr_sig} metrics significant (powerful & conservative)") # Lesson: Pre-register your metrics before running tests!

Practical Code Examples & Templates

End-to-End Classification Evaluation

Python โ€” Complete Classification Pipeline
from sklearn.datasets import load_breast_cancer from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import ( accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, confusion_matrix, classification_report ) # Load data X, y = load_breast_cancer(return_X_y=True) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42, stratify=y ) # Train model = RandomForestClassifier(n_estimators=100, random_state=42) model.fit(X_train, y_train) # Predict y_pred = model.predict(X_test) y_proba = model.predict_proba(X_test)[:, 1] # Evaluate print(f"Accuracy: {accuracy_score(y_test, y_pred):.4f}") print(f"Precision: {precision_score(y_test, y_pred):.4f}") print(f"Recall: {recall_score(y_test, y_pred):.4f}") print(f"F1: {f1_score(y_test, y_pred):.4f}") print(f"AUC-ROC: {roc_auc_score(y_test, y_proba):.4f}") print("\nConfusion Matrix:") print(confusion_matrix(y_test, y_pred)) print("\nDetailed Report:") print(classification_report(y_test, y_pred))

Visual Guide & Decision Trees

Metrics Decision Tree

What's Your Task?
Binary Classification
Multi-Class
Regression
NLP/Generation
Precision, Recall, F1, AUC-ROC, Confusion Matrix
Macro/Weighted F1, Per-Class Metrics
RMSE, MAE, Rยฒ
BLEU, ROUGE, LLM-Judge

Animated Evaluation Pipeline

Complete AI System Evaluation Flow

๐Ÿ“Š
Raw Predictions
โš™๏ธ
Metrics Computation
๐Ÿ”
Disaggregate
๐Ÿ“ˆ
Compare to Baseline
โœ“
Deploy or Investigate

Hands-On Exercises

Exercise 1: Classification Evaluator

Build a Comprehensive Evaluator

Create a function that takes y_true and y_pred, computes accuracy, precision, recall, F1, AUC-ROC. Test on Iris dataset.

Python โ€” Starter Code
def comprehensive_eval(y_true, y_pred, y_proba=None): from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, roc_auc_score results = { 'accuracy': accuracy_score(y_true, y_pred), 'precision': precision_score(y_true, y_pred, average='weighted', zero_division=0), 'recall': recall_score(y_true, y_pred, average='weighted', zero_division=0), 'f1': f1_score(y_true, y_pred, average='weighted', zero_division=0), } if y_proba is not None: results['auc_roc'] = roc_auc_score(y_true, y_proba, average='weighted', multi_class='ovr') return results # TODO: Implement, test on Iris, print results

Exercise 2: A/B Testing

Statistical Significance Test

Given two groups with different success rates, perform t-test, compute p-value and effect size. Determine if difference is significant.

Python โ€” Starter Code
from scipy.stats import ttest_ind import numpy as np # TODO: Generate binary outcomes for two groups # Perform t-test, compute Cohen's d, determine significance

Exercise 3: Fairness Audit

Measure Bias Across Groups

Compute performance metrics separately for different demographic groups. Identify disparities.

Python โ€” Starter Code
import pandas as pd from sklearn.metrics import accuracy_score, f1_score # TODO: Loop over demographic groups, compute metrics, identify disparities

Exercise 4: RAG Evaluation

RAGAS Evaluation

Build a simple RAG system and evaluate with RAGAS metrics.

Python โ€” Starter Code
from ragas import evaluate from ragas.metrics import faithfulness, answer_relevance from datasets import Dataset # TODO: Create test data and run RAGAS evaluation

Interview Questions on AI Evaluation

Q1: Why is accuracy insufficient for imbalanced data? ▼
On imbalanced data (e.g., 99% negative), a model predicting all negatives achieves 99% accuracy but catches 0% of positives. Accuracy assumes equal class importance, which doesn't hold. Use precision, recall, F1, and disaggregate metrics instead.
Q2: Explain precision-recall tradeoff. ▼
Precision = TP/(TP+FP) measures correctness of positive predictions. Recall = TP/(TP+FN) measures fraction of positives caught. They conflict: be conservative (high precision, low recall) or liberal (low precision, high recall). AUC-ROC captures this tradeoff across thresholds.
Q3: BLEU vs ROUGE? ▼
BLEU measures n-gram precision (what fraction of predicted words are in reference). ROUGE measures recall (what fraction of reference words appear in prediction). Use BLEU for translation (word order matters), ROUGE for summarization (content matters).
Q4: How to design A/B test? ▼
Pre-register metrics. Use power analysis to determine sample size. Run at planned duration, don't peek early. Compute p-value and effect size. Use correction for multiple comparisons. Always validate with follow-up test.
Q5: How to detect fairness issues? ▼
Disaggregate metrics by demographic group. Compare TPR, FPR, precision across groups. Measure equalized odds (equal TPR/FPR) or calibration (accurate probabilities in all groups). Document disparities and decide on acceptable tradeoffs.
Q6: Test set vs validation set? ▼
Validation set: used for hyperparameter tuning during training, looked at multiple times. Test set: used ONCE at end for final evaluation, never before. Using test data for tuning causes overfitting to test data.
Q7: How to avoid p-hacking? ▼
Pre-register your metrics before looking at data. Use Bonferroni or FDR correction when testing multiple metrics. If testing 10 metrics at p<0.05, expect 0.5 false positives by chance. Never adjust after seeing results.
Q8: What's a good baseline? ▼
Include: 1) Random (predicts uniformly), 2) Majority class (always predicts most common class), 3) Simple heuristic (domain knowledge), 4) Human performance, 5) Prior published work. Your model must beat all of these.

Frequently Asked Questions

How often should I re-evaluate? ▼
Continuously. Monitor production metrics real-time. Re-evaluate on new data as it arrives. Retrain when performance degrades. Don't evaluate once and forget.
Should I report mean metrics or per-class? ▼
Both. Report overall metrics for summary (accuracy, macro F1). Always disaggregate by class, demographic group, and domain. Most failures hide in disaggregated metrics.
How to handle class imbalance? ▼
1) Stratified splits maintain distribution. 2) Use class-weighted metrics (weighted F1) or per-class metrics. 3) Use threshold-independent metrics (AUC-ROC). 4) Consider cost-sensitive learning if misclassification costs differ.
Statistical significance vs practical significance? ▼
Statistical: p < 0.05 (unlikely due to chance). Practical: difference large enough to matter for business. Tiny improvements can be statistically significant with huge samples but practically irrelevant. Always report effect size (Cohen's d).
How to choose primary metric? ▼
Align with business objective. For fraud detection: prioritize recall (catch fraud). For loan approval: prioritize precision (minimize bad approvals). For search: use engagement metrics. Don't optimize whatever is easiest to measure.
What is overfitting to test data? ▼
Happens when you tune hyperparameters on test data, run many tests until one is 'significant', or select models based on test performance. Makes results look better than production. Always use hold-out test set never looked at before.
How to handle missing/abstaining predictions? ▼
Don't skip examples. Mark as error. Measure abstention rate (fraction of non-predictions). Compute metrics separately on abstaining vs non-abstaining examples. A model that abstains on hard cases might look good but be less useful.
Single test set vs cross-validation? ▼
During development: k-fold CV for robust estimates. For final results: separate hold-out test set opened only once. Ideal: cross-validate during development, report final results on held-out test.
How to evaluate on out-of-distribution data? ▼
Deliberately create OOD test sets. Temporal shift (train 2020, test 2025). Domain shift (train Twitter, test Reddit). Noise (train clean, test noisy). Poor OOD performance signals model won't generalize in production.
Should I optimize for precision or recall? ▼
Depends on cost. For spam detection: optimize precision (avoid false positives = annoyed users). For disease detection: optimize recall (catch all cases = save lives). For fraud: often recall (catch fraud). Document your choice and why.

Summary & Key Takeaways

Core Principles (Revisited)

Measure What Matters

Primary metric = business objective. Avoid optimizing proxies. Accuracy on benchmarks โ‰  user satisfaction in production.

Never Trust One Metric

Accuracy alone hides imbalance and bias. Use multiple metrics: precision+recall+F1+fairness+AUC, disaggregated by group.

Evaluate Continuously

Build evaluation into pipeline. Monitor production real-time. Retrain on distribution shift. One-time evaluation is outdated.

Test on Realistic Data

Test set should match production including edge cases, rare events, distribution shifts. Don't just evaluate on clean benchmarks.

Evaluation Checklist

  1. Define business objective and align primary metric
  2. Create completely held-out test set matching production distribution
  3. Establish baselines: random, heuristic, human, prior work
  4. Train model on training set, tune only on validation set
  5. Evaluate on test set once โ€” don't peek
  6. Compute multiple metrics, disaggregate by group/domain/difficulty
  7. Check for fairness issues explicitly
  8. Monitor production metrics continuously
  9. Retrain on distribution shift or performance degradation
  10. Document all decisions and tradeoffs

Remember: The difference between production AI systems that work and those that fail catastrophically is almost always evaluation rigor, not algorithmic sophistication. Invest in evaluation infrastructure first.

Resources & Further Reading

Essential Python Libraries

  • scikit-learn: Comprehensive metrics via sklearn.metrics
  • RAGAS: RAG evaluation framework (faithfulness, relevance, recall)
  • DeepEval: LLM evaluation with LLM-as-judge and traceability
  • NLTK/ROUGE: NLP metrics (BLEU, ROUGE, linguistic analysis)
  • Weights & Biases: Experiment tracking and metric visualization
  • SciPy (scipy.stats): Statistical testing and power analysis

Key Papers