Fine-Tuning: Adapting Pre-Trained Models
Fine-tuning is the process of taking a pre-trained neural network model and adapting it to perform well on a specific downstream task. Instead of training a model from scratch (which requires massive amounts of data and compute), you leverage the knowledge already captured by the pre-trained model and only adjust its parameters for your specific domain or task.
Pre-trained models like GPT-3, BERT, Mistral, and Llama have learned rich representations of language from enormous amounts of text. Fine-tuning lets you transfer this knowledge to your specific problem โ customer support, medical diagnosis, code generation, or any custom task โ with much less data and compute than training from scratch.
Core Concepts
Transfer Learning
Leverage knowledge learned on large datasets for your specific task, reducing data and compute requirements by 10-100x.
Parameter Efficiency
Modern methods like LoRA fine-tune only a fraction of parameters (0.1-1%), making it practical to run on consumer GPUs.
Task Adaptation
Adapt models for instruction-following, classification, generation, Q&A, and more through specialized tuning approaches.
Alignment
Use RLHF and DPO to align models with human preferences and values, reducing hallucinations and harmful outputs.
Key Difference: Pre-training vs Fine-tuning
Pre-training: Training on massive datasets (billions of tokens) to learn general language patterns. Expensive, done rarely by labs like OpenAI/Anthropic.
Fine-tuning: Training on your specific task/domain dataset (thousands to millions of examples). Cheap, can be done by anyone with a GPU.
Why Fine-Tuning Works
Fine-tuning succeeds because:
- Early layers already know useful features: Early transformer layers have learned to extract phonemes, syntax, named entities, etc. You don't need to relearn these.
- Task-specific information lives in later layers: Later layers learn task-specific patterns. Fine-tuning primarily adjusts these layers.
- Smooth loss landscape: Pre-trained models start near a good solution, so fine-tuning rarely gets stuck in bad local minima.
- Data efficiency: You need maybe 100-1000 examples for fine-tuning vs millions for training from scratch.
Why Fine-Tuning Matters
Fine-tuning is the bridge between cutting-edge research models and practical business applications. It's how companies deploy AI on their specific problems without massive R&D budgets.
The Economics of Fine-Tuning
Training from Scratch (BERT-scale)
Approximate cost/time for adapting a model to your task
Real-World Impact
Customer Support Chatbot
A company fine-tuned Llama-2 on 5,000 internal support tickets and documentation. The fine-tuned model answered customer questions with 92% accuracy, reducing support costs by 40%.
Medical Diagnosis
Researchers fine-tuned PubMedBERT on 10,000 medical notes to predict diagnoses. The fine-tuned model outperformed specialists on rare disease detection by 15%.
Code Generation
A startup fine-tuned CodeLlama on their proprietary codebase and coding standards. The model now generates code that passes tests 78% of the time, vs 35% for the base model.
Methods Comparison: Cost vs Performance
| Method |
GPU Memory |
Training Time |
Performance |
Customization |
| Prompt Engineering |
GPU not needed |
Hours (manual) |
Low-Medium |
Limited |
| Full Fine-Tuning |
24-80 GB |
Hours-Days |
Very High |
Maximal |
| LoRA |
8-16 GB |
30min-4 hours |
High |
Very High |
| QLoRA |
4-8 GB |
1-6 hours |
High |
Very High |
| IA3 |
2-4 GB |
30min-2 hours |
Medium-High |
High |
Key Insight: Full fine-tuning gives maximum performance but is expensive. LoRA gives 95% of full fine-tuning performance at 1/10 the cost. The right method depends on your constraints and performance requirements.
Full Fine-Tuning with HuggingFace Trainer
Full fine-tuning updates all parameters of the model. It's the most flexible approach and gives the best performance, but requires the most compute. Let's build a complete example using the HuggingFace Transformers library.
Setup and Dependencies
pip install torch transformers datasets evaluate scikit-learn
pip install peft accelerate wandb
pip install optuna
Step 1: Load and Prepare Data
from datasets import load_dataset
from transformers import AutoTokenizer
dataset = load_dataset("ag_news")
model_name = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
def preprocess_function(examples):
return tokenizer(
examples["text"],
truncation=True,
max_length=512,
padding="max_length"
)
tokenized_dataset = dataset.map(
preprocess_function,
batched=True,
num_proc=4
)
train_dataset = tokenized_dataset["train"].shuffle(seed=42)
eval_dataset = tokenized_dataset["test"]
print(f"Training samples: {len(train_dataset)}")
print(f"Eval samples: {len(eval_dataset)}")
print(f"Feature keys: {train_dataset.column_names}")
Step 2: Load Pre-Trained Model
from transformers import AutoModelForSequenceClassification
import torch
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
print(f"Using device: {device}")
num_labels = 4
model = AutoModelForSequenceClassification.from_pretrained(
model_name,
num_labels=num_labels,
problem_type="single_label_classification"
).to(device)
print(f"Model parameters: {model.num_parameters():,}")
print(f"Trainable parameters: {sum(p.numel() for p in model.parameters() if p.requires_grad):,}")
Step 3: Configure Training Arguments
from transformers import TrainingArguments
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=3,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
gradient_accumulation_steps=2,
warmup_steps=500,
weight_decay=0.01,
learning_rate=2e-5,
logging_dir="./logs",
logging_steps=100,
evaluation_strategy="steps",
eval_steps=500,
save_strategy="steps",
save_steps=500,
save_total_limit=2,
load_best_model_at_end=True,
metric_for_best_model="accuracy",
greater_is_better=True,
fp16=True,
gradient_checkpointing=True,
report_to=["tensorboard", "wandb"],
)
Step 4: Define Metrics and Trainer
import numpy as np
from datasets import load_metric
from transformers import Trainer, TrainerCallback
metric = load_metric("accuracy")
def compute_metrics(eval_pred):
predictions, labels = eval_pred
predictions = np.argmax(predictions, axis=1)
return metric.compute(predictions=predictions, references=labels)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
compute_metrics=compute_metrics,
callbacks=[],
)
print("Trainer initialized and ready for training")
Step 5: Train and Evaluate
train_result = trainer.train()
eval_results = trainer.evaluate(eval_dataset)
print(f"\nEvaluation Results: {eval_results}")
trainer.save_model("./fine_tuned_model")
print("Model saved to ./fine_tuned_model")
Step 6: Make Predictions
from transformers import pipeline
pipe = pipeline(
"text-classification",
model="./fine_tuned_model",
device=device
)
texts = [
"Apple releases new iPhone with better battery",
"Stock market crashes amid economic concerns",
"Sports: Team wins championship after overtime"
]
predictions = pipe(texts, batch_size=8)
for text, pred in zip(texts, predictions):
print(f"Text: {text}")
print(f" -> Label: {pred['label']}, Score: {pred['score']:.4f}\n")
Key Parameters Explained
Learning Rate (2e-5)
Fine-tuning uses much smaller learning rates than training from scratch (typically 1e-5 to 5e-5). The model already knows a lot โ we're making small adjustments.
Warmup Steps
Gradually increase learning rate at the start. Without warmup, large gradient updates can throw the model off the pre-trained solution.
Weight Decay
L2 regularization penalty. Prevents overfitting when fine-tuning on small datasets.
Gradient Accumulation
Accumulate gradients over multiple batches before updating. Simulates larger batch size without needing more GPU memory.
Memory and Speed Optimizations
- Gradient Checkpointing: Trade compute for memory. Store fewer activations during forward pass, recompute during backward. Reduces memory by ~50%, increases training time by ~20-30%.
- Mixed Precision (fp16): Use 16-bit floats instead of 32-bit. Faster and less memory, nearly identical results.
- Gradient Accumulation: Simulate larger batch sizes without more GPU memory.
- LoRA (next section): Update only a tiny fraction of parameters.
LoRA: Efficient Fine-Tuning with Low-Rank Adaptation
LoRA (Low-Rank Adaptation) is a technique that reduces the number of trainable parameters by 100-1000x while maintaining most of the performance of full fine-tuning. Instead of updating all parameters, LoRA adds small, trainable "adapter" matrices to the model.
The LoRA Idea
The key insight: adaptation to a downstream task doesn't require updating all parameters. Instead, we can express the weight updates as a low-rank decomposition.
LoRA Setup with PEFT
from peft import LoraConfig, get_peft_model
lora_config = LoraConfig(
r=8,
lora_alpha=16,
target_modules=["q_proj", "v_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
Full LoRA Fine-Tuning Example
from transformers import Trainer, TrainingArguments
from peft import LoraConfig, get_peft_model
model = AutoModelForSequenceClassification.from_pretrained(
"bert-base-uncased",
num_labels=4
)
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
lora_config = LoraConfig(
r=8,
lora_alpha=16,
target_modules=["query", "value"],
lora_dropout=0.05,
bias="none",
task_type="SEQ_CLS"
)
model = get_peft_model(model, lora_config)
training_args = TrainingArguments(
output_dir="./lora_results",
num_train_epochs=3,
per_device_train_batch_size=32,
learning_rate=1e-4,
fp16=True,
logging_steps=100,
evaluation_strategy="steps",
eval_steps=500,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
compute_metrics=compute_metrics,
)
trainer.train()
model.save_pretrained("./lora_weights")
print("LoRA weights saved")
LoRA Hyperparameters
| Parameter |
Typical Range |
Effect |
| r (rank) |
4, 8, 16, 32 |
Higher = more capacity but slower. 8 is often optimal. |
| lora_alpha |
8, 16, 32 |
Scaling factor. Usually 2x rank. Controls adapter strength. |
| lora_dropout |
0.05, 0.1, 0.15 |
Regularization. Prevents overfitting on small datasets. |
| target_modules |
["q_proj", "v_proj"] |
Which layers to apply LoRA. Keys have less impact than query/value. |
QLoRA: Even More Efficient
QLoRA combines LoRA with quantization to reduce memory even further. It quantizes the base model to 4-bit, then adds tiny LoRA adapters on top.
from transformers import BitsAndBytesConfig
from peft import LoraConfig, get_peft_model
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_use_double_quant=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Llama-2-7b-hf",
quantization_config=bnb_config,
device_map="auto",
)
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "v_proj", "k_proj", "o_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
print(model.print_trainable_parameters())
LoRA Insight: Most of the fine-tuning updates are low-rank! This is why LoRA works so well. The model is doing something fundamentally different for your task, but that difference lives in a low-rank subspace.
Loading and Using LoRA Models
from peft import AutoPeftModelForSequenceClassification
model = AutoPeftModelForSequenceClassification.from_pretrained(
"path/to/lora_weights"
)
outputs = model(input_ids, attention_mask=attention_mask)
merged_model = model.merge_and_unload()
merged_model.save_pretrained("./merged_model")
Instruction Tuning: Teaching Models to Follow Instructions
Instruction tuning teaches language models to follow user instructions and respond appropriately. Instead of just predicting the next token, the model learns to understand tasks described in natural language and execute them correctly.
Why Instruction Tuning?
Pre-trained language models are trained to predict the next token. They don't inherently understand task descriptions or follow instructions. A model might complete "The capital of France is" with accurate information, but fail at "Answer in French: What is the capital of France?"
Instruction tuning aligns the model with human expectations: given a task description (prompt), generate the correct output.
Dataset Format
{
"instruction": "Classify the sentiment of this review.",
"input": "This movie was absolutely terrible. Boring, slow, and a waste of time.",
"output": "negative"
}
{
"instruction": "Summarize the following article in 2-3 sentences.",
"input": "New research shows that regular exercise...",
"output": "The study found that..."
}
{
"instruction": "Translate to French",
"input": "Hello, how are you?",
"output": "Bonjour, comment allez-vous?"
}
{
"instruction": "Write a Python function that...",
"input": "that takes a list and returns unique elements",
"output": "def unique(lst):\n return list(set(lst))"
}
Building an Instruction Dataset
import json
from datasets import Dataset
examples = [
{
"instruction": "Classify the sentiment",
"input": "This product is amazing!",
"output": "positive"
},
{
"instruction": "Classify the sentiment",
"input": "Worst purchase ever",
"output": "negative"
},
]
def format_instruction(example):
return {
"text": f"""Instruction: {example['instruction']}
Input: {example['input']}
Output: {example['output']}<|endoftext|>"""
}
formatted = [format_instruction(ex) for ex in examples]
dataset = Dataset.from_dict({
"text": [ex["text"] for ex in formatted]
})
print(f"Dataset size: {len(dataset)} examples")
Fine-Tuning for Instruction Following
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
Trainer,
TrainingArguments,
)
model_name = "gpt2"
model = AutoModelForCausalLM.from_pretrained(model_name)
tokenizer = AutoTokenizer.from_pretrained(model_name)
def preprocess_function(examples):
return tokenizer(
examples["text"],
truncation=True,
max_length=512,
)
train_dataset = dataset.map(preprocess_function, batched=True)
args = TrainingArguments(
output_dir="./instruction_model",
num_train_epochs=3,
per_device_train_batch_size=8,
learning_rate=5e-5,
warmup_steps=100,
weight_decay=0.01,
logging_steps=10,
save_strategy="epoch",
fp16=True,
)
trainer = Trainer(
model=model,
args=args,
train_dataset=train_dataset,
)
trainer.train()
model.save_pretrained("./instruction_tuned_model")
Inference: Using the Instruction-Tuned Model
from transformers import pipeline
pipe = pipeline(
"text-generation",
model="./instruction_tuned_model",
device=0
)
prompt = """Instruction: Classify sentiment
Input: This is great!
Output:"""
output = pipe(
prompt,
max_length=50,
num_return_sequences=1,
temperature=0.7,
top_p=0.9,
)
print(output[0]["generated_text"])
Best Practices for Instruction Data
Diversity
Include variety: different tasks, domains, difficulty levels. A dataset with only classification gets overtrained on that task.
Quality > Quantity
100 high-quality instruction examples beat 10,000 low-quality ones. Spend time curating.
Clear Instructions
Write instructions that a human could understand. Avoid ambiguity. Bad: 'Do something'. Good: 'Classify the sentiment as positive, negative, or neutral.'
Balanced Outputs
If possible, balance output lengths and task types. Prevents bias toward one type of response.
Instruction Tuning vs Full Fine-Tuning
| Aspect |
Instruction Tuning |
Full Fine-Tuning |
| Task Type |
General purpose (any task) |
Specific task |
| Data Format |
instruction + input โ output |
input โ output |
| Model Flexibility |
High (understands new tasks) |
Low (locked to task) |
| Data Required |
500-5000 diverse examples |
100-1000 task-specific examples |
| Best For |
Building helpful assistants |
Specific applications |
Alignment: RLHF and Direct Preference Optimization (DPO)
After instruction tuning, models often hallucinate, refuse valid requests, or produce harmful content. RLHF (Reinforcement Learning from Human Feedback) and DPO (Direct Preference Optimization) align models with human preferences and values.
The Problem: Why Fine-Tuning Isn't Enough
Consider these issues:
- Hallucination: Model confidently states false information.
- Refusal: Model refuses to help with legitimate requests.
- Helpfulness: Model gives technically correct but unhelpful answers.
- Harmfulness: Model generates toxic, biased, or dangerous content.
Standard fine-tuning on examples doesn't address these โ we need to optimize for preferences.
RLHF: Reinforcement Learning from Human Feedback
RLHF has 4 phases:
1. Supervised Fine-Tuning
Fine-tune on high-quality instruction examples (previous section).
2. Reward Modeling
Collect human preferences (good vs bad outputs). Train a model to predict which output humans prefer.
3. Reinforcement Learning
Use the reward model to train the language model. Maximize expected reward while staying close to pre-trained model.
4. Red Teaming
Find and fix failure modes. Repeat until model is safe and helpful.
Reward Model Training
from transformers import AutoModelForSequenceClassification, Trainer
import torch
preference_data = [
{
"prompt": "Tell me a joke",
"chosen": "Why did the chicken cross the road? To get to the other side!",
"rejected": "I cannot tell jokes."
},
]
class PairwiseDataset(Dataset):
def __init__(self, data, tokenizer):
self.tokenizer = tokenizer
self.data = data
def __len__(self):
return len(self.data)
def __getitem__(self, idx):
example = self.data[idx]
chosen = self.tokenizer(
example["prompt"] + example["chosen"],
max_length=512,
truncation=True,
return_tensors="pt"
)
rejected = self.tokenizer(
example["prompt"] + example["rejected"],
max_length=512,
truncation=True,
return_tensors="pt"
)
return {
"chosen_input_ids": chosen["input_ids"].squeeze(),
"chosen_attention_mask": chosen["attention_mask"].squeeze(),
"rejected_input_ids": rejected["input_ids"].squeeze(),
"rejected_attention_mask": rejected["attention_mask"].squeeze(),
}
model = AutoModelForSequenceClassification.from_pretrained(
"gpt2",
num_labels=1
)
def reward_loss(model_output_chosen, model_output_rejected):
return -torch.nn.functional.logsigmoid(
model_output_chosen.logits - model_output_rejected.logits
).mean()
print("Reward model training setup ready")
Direct Preference Optimization (DPO)
DPO is a simpler alternative to RLHF that doesn't require training a separate reward model. Instead, it directly optimizes the language model using preference data.
DPO Implementation
from trl import DPOTrainer, DPOConfig
from transformers import AutoModelForCausalLM, AutoTokenizer
preference_dataset = [
{
"prompt": "Tell me a joke",
"chosen": "Why did the chicken cross the road?",
"rejected": "I can't tell jokes.",
},
]
model = AutoModelForCausalLM.from_pretrained("gpt2")
tokenizer = AutoTokenizer.from_pretrained("gpt2")
dpo_config = DPOConfig(
beta=0.1,
learning_rate=1e-5,
num_train_epochs=3,
per_device_train_batch_size=4,
output_dir="./dpo_model",
)
trainer = DPOTrainer(
model=model,
args=dpo_config,
train_dataset=preference_dataset,
tokenizer=tokenizer,
)
trainer.train()
model.save_pretrained("./dpo_model")
RLHF vs DPO Comparison
| Aspect |
RLHF |
DPO |
| Complexity |
High (4 phases) |
Low (direct optimization) |
| Training Cost |
Expensive (RL + reward model) |
Cheap (direct training) |
| Reward Model |
Separate model needed |
No separate model |
| Stability |
Can be unstable |
More stable |
| Performance |
Excellent |
Excellent |
| When to Use |
Production systems needing fine control |
Most practitioners |
Key Insight: DPO has become increasingly popular because it's simpler than RLHF while achieving similar or better results. Unless you have specific reasons for RLHF, start with DPO.
Dataset Preparation for Fine-Tuning
Quality data is everything in fine-tuning. A small dataset of high-quality examples beats a large dataset of noisy examples. Let's cover practical dataset preparation techniques.
Data Sources and Collection
Internal Data
Customer tickets, documentation, past conversations. Most valuable! Domain-specific and private.
Public Datasets
HuggingFace Hub, Kaggle, academic datasets. Free but may not match your domain exactly.
Synthetic Data
Generate examples using GPT-4 or other strong models. Useful for augmentation or when real data is scarce.
Crowdsourcing
Hire annotators to create or label data. Expensive but high quality if managed well.
Data Cleaning and Preprocessing
import json
import re
from datasets import Dataset
def clean_text(text):
'''Remove noise from text'''
text = re.sub(r'<[^>]+>', '', text)
text = re.sub(r'\s+', ' ', text).strip()
text = re.sub(r'[^a-zA-Z0-9\s.!?,-]', '', text)
return text
def validate_example(example):
'''Check if example is valid for fine-tuning'''
if len(example.get("input", "")) < 5:
return False
if len(example.get("output", "")) < 5:
return False
if "input" not in example or "output" not in example:
return False
return True
with open("raw_data.jsonl") as f:
raw_examples = [json.loads(line) for line in f]
cleaned = []
for ex in raw_examples:
if validate_example(ex):
ex["input"] = clean_text(ex["input"])
ex["output"] = clean_text(ex["output"])
cleaned.append(ex)
print(f"Kept {len(cleaned)}/{len(raw_examples)} examples")
dataset = Dataset.from_dict({
"input": [ex["input"] for ex in cleaned],
"output": [ex["output"] for ex in cleaned],
})
Train/Validation/Test Split
from sklearn.model_selection import train_test_split
train, temp = train_test_split(cleaned, test_size=0.2, random_state=42)
val, test = train_test_split(temp, test_size=0.5, random_state=42)
print(f"Train: {len(train)} | Val: {len(val)} | Test: {len(test)}")
train_dataset = Dataset.from_dict({
"input": [ex["input"] for ex in train],
"output": [ex["output"] for ex in train],
})
val_dataset = Dataset.from_dict({
"input": [ex["input"] for ex in val],
"output": [ex["output"] for ex in val],
})
test_dataset = Dataset.from_dict({
"input": [ex["input"] for ex in test],
"output": [ex["output"] for ex in test],
})
train_dataset.save_to_disk("./train_data")
val_dataset.save_to_disk("./val_data")
test_dataset.save_to_disk("./test_data")
Data Augmentation Techniques
import random
def paraphrase_instruction(instruction):
'''Simple paraphrasing of instructions'''
templates = [
"Can you {}?",
"Please {}.",
"I need you to {}.",
"Would you {}?",
"How would you {}?",
]
base = instruction.lower()
template = random.choice(templates)
return template.format(base)
def add_variations(dataset):
'''Augment dataset with variations'''
augmented = []
for example in dataset:
augmented.append(example)
if "instruction" in example:
aug_ex = example.copy()
aug_ex["instruction"] = paraphrase_instruction(example["instruction"])
augmented.append(aug_ex)
return augmented
augmented_train = add_variations(train)
print(f"Training size after augmentation: {len(augmented_train)}")
Data Analysis and Debugging
import numpy as np
from collections import Counter
input_lengths = [len(ex["input"].split()) for ex in dataset]
output_lengths = [len(ex["output"].split()) for ex in dataset]
print(f"Input length - Mean: {np.mean(input_lengths):.0f}, "
f"Min: {np.min(input_lengths)}, Max: {np.max(input_lengths)}")
print(f"Output length - Mean: {np.mean(output_lengths):.0f}, "
f"Min: {np.min(output_lengths)}, Max: {np.max(output_lengths)}")
inputs = [ex["input"] for ex in dataset]
print(f"Duplicates: {len(inputs) - len(set(inputs))}")
if "label" in dataset[0]:
labels = [ex["label"] for ex in dataset]
print("Label distribution:", Counter(labels))
Dataset Size Guidelines
| Fine-Tuning Method |
Minimum Data |
Recommended |
Optimal |
| Full Fine-Tuning (Classification) |
100 |
500-1000 |
5000+ |
| Full Fine-Tuning (Generation) |
200 |
1000-2000 |
10000+ |
| LoRA (Classification) |
50 |
200-500 |
2000+ |
| LoRA (Generation) |
100 |
500-1000 |
5000+ |
| Instruction Tuning |
200 |
1000-2000 |
5000+ |
Rule of Thumb: You need at least 100 examples to improve over a pre-trained model. 500-2000 gives good results for most tasks. More data always helps, but with diminishing returns.
Evaluation During Training: Metrics and Monitoring
You can't improve what you don't measure. Proper evaluation during fine-tuning is crucial for understanding progress and preventing overfitting.
Classification Metrics
from sklearn.metrics import (
accuracy_score,
precision_recall_fscore_support,
confusion_matrix,
roc_auc_score,
classification_report,
)
import numpy as np
def compute_classification_metrics(predictions, labels, num_classes=None):
'''Compute standard classification metrics'''
accuracy = accuracy_score(labels, predictions)
precision, recall, f1, _ = precision_recall_fscore_support(
labels, predictions, average="weighted"
)
report = classification_report(labels, predictions, output_dict=True)
cm = confusion_matrix(labels, predictions)
return {
"accuracy": accuracy,
"precision": precision,
"recall": recall,
"f1": f1,
"per_class": report,
"confusion_matrix": cm,
}
predictions = np.argmax(logits, axis=1)
metrics = compute_classification_metrics(predictions, labels)
print(f"Accuracy: {metrics['accuracy']:.4f}")
print(f"F1: {metrics['f1']:.4f}")
Generation Metrics
from rouge_score import rouge_scorer
from nltk.translate.bleu_score import corpus_bleu
import nltk
nltk.download('punkt')
def compute_generation_metrics(predictions, references):
'''Compute metrics for generation tasks (summarization, translation)'''
rouge = rouge_scorer.RougeScorer(['rouge1', 'rougeL'], use_stemmer=True)
rouge_scores = [rouge.score(ref, pred) for pred, ref in zip(predictions, references)]
avg_rouge1 = np.mean([s['rouge1'].fmeasure for s in rouge_scores])
references_tokens = [[ref.split()] for ref in references]
predictions_tokens = [pred.split() for pred in predictions]
bleu = corpus_bleu(references_tokens, predictions_tokens)
exact_match = sum([p.strip() == r.strip() for p, r in zip(predictions, references)]) / len(predictions)
return {
"rouge1": avg_rouge1,
"bleu": bleu,
"exact_match": exact_match,
}
preds = ["The cat sat on the mat", "Dogs are animals"]
refs = ["A cat was sitting on the mat", "Dogs are mammals"]
metrics = compute_generation_metrics(preds, refs)
print(f"ROUGE-1: {metrics['rouge1']:.4f}")
Custom Evaluation Metrics
def evaluate_code_generation(predictions, references):
'''Custom metric: code generation accuracy'''
scores = []
for pred, ref in zip(predictions, references):
try:
pred_result = eval(pred)
ref_result = eval(ref)
scores.append(1.0 if pred_result == ref_result else 0.0)
except:
scores.append(0.0)
return np.mean(scores)
def evaluate_medical_diagnosis(predictions, references, task="sensitivity"):
'''Custom metric: sensitivity/specificity trade-off'''
from sklearn.metrics import recall_score, precision_score
if task == "sensitivity":
return recall_score(references, predictions)
elif task == "specificity":
return recall_score(references, predictions, pos_label=0)
def compute_metrics(eval_pred):
predictions, labels = eval_pred
predictions = np.argmax(predictions, axis=1)
accuracy = accuracy_score(labels, predictions)
sensitivity = evaluate_medical_diagnosis(predictions, labels, "sensitivity")
return {
"accuracy": accuracy,
"sensitivity": sensitivity,
}
Monitoring Training with Logging
import wandb
from transformers import TrainingArguments, Trainer
wandb.init(project="fine-tuning", name="my-experiment")
training_args = TrainingArguments(
output_dir="./results",
num_train_epochs=3,
per_device_train_batch_size=16,
logging_dir="./logs",
logging_steps=100,
evaluation_strategy="steps",
eval_steps=500,
report_to=["wandb"],
run_name="gpt2-finetuning-v1",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
compute_metrics=compute_metrics,
callbacks=[],
)
wandb.log({
"dataset_size": len(train_dataset),
"model_parameters": model.num_parameters(),
})
trainer.train()
Detecting Overfitting
Look for these signs during training:
- Training loss keeps decreasing: Good, model is learning.
- Validation loss increases: Overfitting! Regularize more (higher dropout, weight decay, data augmentation).
- Training and validation loss diverge: Model overfitting. Too many epochs or not enough data.
- Metrics plateau: Diminishing returns. Training longer won't help. Try more data, adjust hyperparameters, or accept current performance.
Early Stopping
Stop training when validation loss stops improving for N evaluations. Prevents overfitting and saves compute. Most trainers support this via callbacks.
Hyperparameter Search and Optimization
Fine-tuning success depends heavily on hyperparameters: learning rate, batch size, warmup steps, weight decay, etc. Random hyperparameter search is inefficient. Optuna enables systematic optimization.
Common Hyperparameters to Tune
| Hyperparameter |
Typical Range |
Impact |
Tuning Order |
| Learning Rate |
1e-5 to 5e-4 |
Very High |
1st |
| Batch Size |
8 to 128 |
High |
2nd |
| Warmup Steps |
0 to 1000 |
Medium |
3rd |
| Weight Decay |
0 to 0.1 |
Medium |
3rd |
| Epochs |
2 to 10 |
Medium |
4th |
Optuna-Based Hyperparameter Search
import optuna
from transformers import Trainer, TrainingArguments
import numpy as np
def objective(trial):
'''Objective function for Optuna to optimize'''
learning_rate = trial.suggest_float("learning_rate", 1e-5, 5e-4, log=True)
batch_size = trial.suggest_int("batch_size", 8, 32)
warmup_steps = trial.suggest_int("warmup_steps", 0, 1000, step=100)
weight_decay = trial.suggest_float("weight_decay", 0.0, 0.1)
training_args = TrainingArguments(
output_dir=f"./trial_{trial.number}",
num_train_epochs=3,
per_device_train_batch_size=batch_size,
per_device_eval_batch_size=32,
learning_rate=learning_rate,
warmup_steps=warmup_steps,
weight_decay=weight_decay,
evaluation_strategy="steps",
eval_steps=500,
save_strategy="no",
logging_steps=100,
fp16=True,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=train_dataset,
eval_dataset=eval_dataset,
compute_metrics=compute_metrics,
)
trainer.train()
eval_result = trainer.evaluate()
return eval_result["eval_accuracy"]
study = optuna.create_study(direction="maximize")
study.optimize(objective, n_trials=20)
best_trial = study.best_trial
print(f"Best accuracy: {best_trial.value:.4f}")
print(f"Best hyperparameters: {best_trial.params}")
Grid Search vs Random Search vs Bayesian Optimization
| Method |
Efficiency |
Complexity |
Best For |
| Grid Search |
Low |
Low |
Small search spaces |
| Random Search |
Medium |
Low |
Quick experiments |
| Bayesian Opt (Optuna) |
High |
Medium |
Production tuning |
Learning Rate Scheduling
from transformers import TrainingArguments
import math
training_args = TrainingArguments(
output_dir="./results",
learning_rate=2e-5,
warmup_steps=500,
lr_scheduler_type="linear",
lr_scheduler_type="cosine",
lr_scheduler_type="cosine_with_restarts",
num_cycles=3,
)
class CustomScheduleCallback:
def on_step_end(self, args, state, control, **kwargs):
pass
Learning Rate Finder
import matplotlib.pyplot as plt
import numpy as np
def learning_rate_finder(model, train_loader, device, start_lr=1e-6, end_lr=1e-1, num_iterations=100):
'''Find optimal learning rate by training with increasing LR'''
lrs = np.logspace(np.log10(start_lr), np.log10(end_lr), num_iterations)
losses = []
for lr in lrs:
for param_group in optimizer.param_groups:
param_group['lr'] = lr
model.train()
batch = next(iter(train_loader))
outputs = model(batch)
loss = outputs.loss
optimizer.zero_grad()
loss.backward()
optimizer.step()
losses.append(loss.item())
plt.figure(figsize=(10, 6))
plt.xscale('log')
plt.plot(lrs, losses)
plt.xlabel('Learning Rate')
plt.ylabel('Loss')
plt.title('Learning Rate Finder')
plt.savefig('./lr_finder.png')
best_lr = lrs[np.argmin(losses)]
print(f"Suggested learning rate: {best_lr:.2e}")
return best_lr
Comprehensive Comparison: Full vs LoRA vs QLoRA vs Adapters
Multiple efficient fine-tuning methods exist. Each has trade-offs in memory, speed, and performance. This section compares them across practical dimensions.
Visual Comparison: Memory vs Performance
Full Fine-Tuning
100% Performance, 100% Memory
LoRA (r=8)
95% Performance, 30% Memory
QLoRA
92% Performance, 10% Memory
IA3
85% Performance, 5% Memory
Prompt Tuning
70% Performance, 2% Memory
Approximate relative performance and memory usage for a 7B model
Detailed Comparison Table
| Method |
Trainable Params |
GPU Memory |
Training Time |
Final Performance |
Complexity |
| Full Fine-Tuning |
100% |
80 GB (7B model) |
8-24h |
100% |
Low |
| LoRA |
0.1-1% |
24-40 GB |
2-8h |
95-98% |
Low |
| QLoRA |
0.1-1% |
8-16 GB |
4-12h |
92-96% |
Medium |
| IA3 |
0.01-0.1% |
4-8 GB |
2-6h |
80-90% |
Low |
| Prefix Tuning |
0.1-1% |
16-24 GB |
3-6h |
90-95% |
Medium |
| Prompt Tuning |
0.01-0.1% |
2-4 GB |
30min-2h |
70-85% |
Low |
Which Method to Use?
Full Fine-Tuning
Use when: Maximum performance is critical, you have compute budget, task is very different from pre-training. Gives best results.
LoRA
Use when: Good performance with moderate efficiency needed. Most practitioners should start here. Great balance of quality and cost.
QLoRA
Use when: Limited GPU memory (RTX 3090, RTX 4090). Adds 20-30% training time but saves 5x memory vs LoRA.
IA3
Use when: Extreme memory constraints or many small adaptations needed. Performance trade-off is acceptable.
Combining Methods: Multi-Adapter Fine-Tuning
Advanced: Train multiple task-specific adapters and switch between them at inference time.
from peft import PeftModel, LoraConfig, get_peft_model
base_model = AutoModelForCausalLM.from_pretrained("mistral-7b")
adapters = {}
tasks = ["sentiment", "translation", "summarization"]
for task in tasks:
lora_config = LoraConfig(
r=8,
lora_alpha=16,
target_modules=["q_proj", "v_proj"],
task_type="CAUSAL_LM"
)
model = get_peft_model(base_model, lora_config)
trainer.train()
model.save_pretrained(f"./adapters/{task}")
adapters[task] = model
loaded_model = PeftModel.from_pretrained(
base_model,
f"./adapters/sentiment"
)
loaded_model.set_adapter("translation")
Practical Tip: If you're unsure, start with LoRA. It has great performance, is memory efficient, and is well-supported across frameworks.
Fine-Tuning Best Practices and Common Pitfalls
10 Essential Best Practices
1. Start Small and Iterate
Don't fine-tune on 10,000 examples immediately. Start with 100-200 examples. If performance is poor, the issue is likely not data size but hyperparameters or data quality.
2. Always Use a Validation Set
Monitor validation metrics during training. Without validation, you can't detect overfitting. Use early stopping to prevent wasted compute.
3. Lower Learning Rate Than Pre-training
Fine-tuning uses 10-100x smaller learning rates than pre-training. Typically 1e-5 to 5e-4. Too high and you destroy pre-trained knowledge. Too low and training is painfully slow.
4. Use Mixed Precision Training
Enable fp16=True in training args. Speeds up training by 20-30%, uses half the memory, and accuracy is nearly identical. It's a free win.
5. Data Quality Beats Quantity
100 perfect examples beat 10,000 noisy ones. Spend time cleaning, deduplicating, and validating data. This is often the highest ROI task.
6. Try Different Model Sizes
Bigger models are often better, but smaller models can work for simple tasks. A 7B model fine-tuned on your data might beat a 70B base model with prompts.
7. Monitor for Mode Collapse
If your fine-tuned model generates identical outputs regardless of input, it's collapsed. Increase data diversity, reduce regularization, or lower learning rate.
8. Save Checkpoints Regularly
Don't save only the final model. Save intermediate checkpoints. Sometimes earlier checkpoints are better due to overfitting.
9. Use Gradient Accumulation for Large Batches
If you can't fit large batches in memory, use gradient accumulation. Simulates larger batches with equivalent results but lower memory.
10. Test on Domain-Specific Data
Evaluate on examples from your actual use case, not just public benchmarks. A model that scores 95% on GLUE might perform poorly on your specific task.
Common Pitfalls and Solutions
Likely causes:
- Learning rate too high โ reduces to 1e-5
- Not enough training data โ gather more examples
- Data distribution mismatch โ ensure data is from your domain
- Task is too different from pre-training โ try different architecture
- Reduce batch size (try 2-4 instead of 32)
- Enable gradient checkpointing: gradient_checkpointing=True
- Use LoRA instead of full fine-tuning
- Use QLoRA for even more memory savings
- Reduce max_seq_length (e.g., 256 instead of 512)
- Increase weight decay (try 0.01 or 0.1)
- Increase dropout (lora_dropout=0.1 or higher)
- Reduce number of epochs (try 1-2 instead of 3-5)
- Add data augmentation
- Use early stopping based on validation loss
Advanced Fine-Tuning Techniques
Multi-task Fine-Tuning
Train on multiple tasks simultaneously. The model learns shared representations and often performs better on individual tasks.
from torch.utils.data import ConcatDataset, DataLoader
sentiment_data = load_dataset("amazon_reviews_multi")["train"]
topic_data = load_dataset("ag_news")["train"]
def add_task_marker(example, task_name):
example["task"] = task_name
example["text"] = f"Task: {task_name}\n{example['text']}"
return example
sentiment_data = sentiment_data.map(
lambda x: add_task_marker(x, "sentiment")
)
topic_data = topic_data.map(
lambda x: add_task_marker(x, "topic")
)
combined = ConcatDataset([sentiment_data, topic_data])
trainer = Trainer(...)
trainer.train()
Continued Pre-training (Domain Adaptation)
Sometimes better than fine-tuning: continue pre-training on unlabeled data from your domain before fine-tuning on labeled data.
training_args = TrainingArguments(
output_dir="./domain_adapted_model",
num_train_epochs=1,
per_device_train_batch_size=16,
learning_rate=5e-5,
fp16=True,
)
unlabeled_data = load_dataset("text", data_files="domain_corpus.txt")
trainer = Trainer(
model=model,
args=training_args,
train_dataset=unlabeled_data,
data_collator=DataCollatorForLanguageModeling(tokenizer, mlm=False),
)
trainer.train()
model = AutoModelForSequenceClassification.from_pretrained(
"./domain_adapted_model",
num_labels=4
)
trainer = Trainer(...)
trainer.train()
Few-Shot In-Context Learning
Alternative to fine-tuning: include examples in the prompt.
from transformers import pipeline
pipe = pipeline("text-classification", model="gpt2")
prompt = """Classify the sentiment.
Example 1:
Input: Great movie!
Output: positive
Example 2:
Input: Terrible experience
Output: negative
Now classify:
Input: This is okay.
Output:"""
output = pipe(prompt, max_length=200)
print(output)
Knowledge Distillation
Train a smaller, faster model to mimic a larger fine-tuned model.
import torch
import torch.nn.functional as F
def distillation_loss(student_logits, teacher_logits, labels, temperature=3.0, alpha=0.7):
'''Combined supervised + distillation loss'''
ce_loss = F.cross_entropy(student_logits, labels)
kl_loss = F.kl_div(
F.log_softmax(student_logits / temperature, dim=1),
F.softmax(teacher_logits / temperature, dim=1),
reduction='batchmean'
) * (temperature ** 2)
loss = alpha * ce_loss + (1 - alpha) * kl_loss
return loss
for batch in train_loader:
student_logits = student_model(batch)
with torch.no_grad():
teacher_logits = teacher_model(batch)
loss = distillation_loss(student_logits, teacher_logits, batch['labels'])
loss.backward()
optimizer.step()
Deployment: Saving, Merging, and Serving Fine-Tuned Models
Saving Fine-Tuned Models
from peft import AutoPeftModelForCausalLM
model.save_pretrained("./full_ft_model")
peft_model.save_pretrained("./lora_weights")
tokenizer.save_pretrained("./full_ft_model")
tokenizer.save_pretrained("./lora_weights")
Merging LoRA Adapters with Base Model
For deployment, you may want to merge LoRA weights into the base model. This creates a single model file without needing PEFT at inference.
from peft import PeftModel, PeftConfig
base_model = AutoModelForCausalLM.from_pretrained("mistral-7b")
peft_model = PeftModel.from_pretrained(base_model, "./lora_weights")
merged_model = peft_model.merge_and_unload()
merged_model.save_pretrained("./merged_model")
tokenizer.save_pretrained("./merged_model")
from transformers import pipeline
pipe = pipeline("text-generation", model="./merged_model")
Serving Fine-Tuned Models
from fastapi import FastAPI
from transformers import pipeline
import torch
app = FastAPI()
device = 0 if torch.cuda.is_available() else -1
pipe = pipeline(
"text-generation",
model="./fine_tuned_model",
device=device,
torch_dtype=torch.float16,
)
@app.post("/predict")
def predict(text: str, max_length: int = 100):
outputs = pipe(text, max_length=max_length, num_return_sequences=1)
return {"output": outputs[0]["generated_text"]}
Quantization for Efficient Serving
from optimum.onnxruntime import ORTModelForSequenceClassification
ort_model = ORTModelForSequenceClassification.from_pretrained(
"./fine_tuned_model",
from_transformers=True,
file_name="model_quantized.onnx"
)
ort_model.save_pretrained("./quantized_model")
from transformers import pipeline
pipe = pipeline("text-classification", model="./quantized_model")
Model Comparison and Versioning
Keep track of multiple fine-tuned versions:
import json
from datetime import datetime
model_registry = {
"v1": {
"date": "2024-01-15",
"method": "full-ft",
"dataset_size": 5000,
"val_accuracy": 0.92,
"path": "./models/ft_v1",
},
"v2": {
"date": "2024-01-20",
"method": "lora",
"dataset_size": 10000,
"val_accuracy": 0.94,
"path": "./models/ft_v2",
}
}
with open("model_registry.json", "w") as f:
json.dump(model_registry, f, indent=2)
best_model_path = model_registry["v2"]["path"]
Troubleshooting Common Issues
Training Issues
Solutions (in order of effectiveness):
- Reduce batch size:
per_device_train_batch_size=4 (from 16)
- Enable gradient checkpointing:
gradient_checkpointing=True
- Reduce max sequence length:
max_length=256 (from 512)
- Enable 8-bit optimization: Install bitsandbytes, use
load_in_8bit=True
- Use LoRA: reduces memory by 50-70%
- Use QLoRA: reduces memory by 90%+
- NaN loss: Usually learning rate too high. Reduce by 10x.
- Loss stays 0: Model has converged or collapsed. Try different initialization or check data.
- Solutions: Use smaller learning rate, add warmup steps, check data for NaN values, try different random seed.
- Mode collapse: Model learned a single "safe" output.
- Solutions: Increase data diversity, lower learning rate, add temperature to decoding (temperature=0.7), check training examples for quality.
Data Issues
- Use Arrow format (.arrow) instead of JSON for faster loading
- Set
num_proc=4 in dataset.map() for parallel processing
- Cache preprocessed data:
dataset.save_to_disk("./cached_data")
- Use smaller batches to see progress faster during development
- Check what special tokens your tokenizer has:
tokenizer.special_tokens_map
- Add custom tokens:
tokenizer.add_special_tokens({'additional_special_tokens': ['', '']})
- Resize model embeddings:
model.resize_token_embeddings(len(tokenizer))
Inference Issues
- Check that you're loading the right model (fine-tuned version, not base)
- Evaluate on task-specific metrics, not generic ones
- Fine-tune longer: increase epochs to 5-10
- Add more data: may need 500+ examples for measurable improvement
- Check data quality: bad labels hurt more than no labels
Complete End-to-End Code Examples
Example 1: Fine-Tune BERT for Sentiment Classification
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSequenceClassification
from transformers import TrainingArguments, Trainer
import torch
dataset = load_dataset("rotten_tomatoes")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
def tokenize_function(examples):
return tokenizer(examples["text"], padding="max_length", truncation=True, max_length=128)
tokenized = dataset.map(tokenize_function, batched=True)
model = AutoModelForSequenceClassification.from_pretrained(
"bert-base-uncased",
num_labels=2
)
training_args = TrainingArguments(
output_dir="./bert_sentiment",
num_train_epochs=3,
per_device_train_batch_size=32,
per_device_eval_batch_size=64,
learning_rate=2e-5,
logging_steps=100,
evaluation_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="accuracy",
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
)
trainer.train()
from transformers import pipeline
pipe = pipeline("text-classification", model="./bert_sentiment")
print(pipe("This movie is amazing!"))
Example 2: LoRA Fine-Tune Mistral 7B
from peft import LoraConfig, get_peft_model
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
Trainer,
TrainingArguments,
BitsAndBytesConfig,
)
import torch
bnb_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
"mistralai/Mistral-7B-v0.1",
quantization_config=bnb_config,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.1")
lora_config = LoraConfig(
r=16,
lora_alpha=32,
target_modules=["q_proj", "v_proj"],
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM"
)
model = get_peft_model(model, lora_config)
print(model.print_trainable_parameters())
dataset = load_dataset("openwebtext", split="train[:1%]")
def tokenize(examples):
return tokenizer(examples["text"], truncation=True, max_length=512)
tokenized = dataset.map(tokenize, batched=True, remove_columns=["text"])
training_args = TrainingArguments(
output_dir="./mistral_lora",
num_train_epochs=1,
per_device_train_batch_size=4,
gradient_accumulation_steps=4,
learning_rate=5e-5,
warmup_steps=100,
logging_steps=10,
save_steps=500,
fp16=True,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized,
)
trainer.train()
model.save_pretrained("./mistral_lora")
Example 3: Instruction Tuning GPT-2
import json
from datasets import Dataset
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
Trainer,
TrainingArguments,
DataCollatorForLanguageModeling,
)
instructions = [
{
"instruction": "Classify the sentiment",
"input": "This product is great!",
"output": "positive"
},
{
"instruction": "Classify the sentiment",
"input": "Terrible experience",
"output": "negative"
},
]
def format_example(example):
text = f"Instruction: {example['instruction']}\n"
text += f"Input: {example['input']}\n"
text += f"Output: {example['output']}<|endoftext|>"
return {"text": text}
formatted = [format_example(ex) for ex in instructions]
dataset = Dataset.from_dict({"text": [ex["text"] for ex in formatted]})
tokenizer = AutoTokenizer.from_pretrained("gpt2")
tokenizer.pad_token = tokenizer.eos_token
def tokenize(examples):
return tokenizer(examples["text"], truncation=True, max_length=512)
tokenized = dataset.map(tokenize, batched=True)
model = AutoModelForCausalLM.from_pretrained("gpt2")
training_args = TrainingArguments(
output_dir="./instruction_gpt2",
num_train_epochs=3,
per_device_train_batch_size=8,
learning_rate=5e-5,
warmup_steps=100,
save_steps=100,
save_total_limit=2,
)
trainer = Trainer(
model=model,
args=training_args,
train_dataset=tokenized,
data_collator=DataCollatorForLanguageModeling(tokenizer, mlm=False),
)
trainer.train()
Hands-On Exercises
Exercise 1: Fine-Tune a Model on Your Own Data
Task: Collect or create 100-200 examples of text with labels (sentiment, intent, topic, etc.). Fine-tune a BERT or RoBERTa model on your data.
Learning goals: Understand the full fine-tuning pipeline, learn how to preprocess real data, evaluate on test set.
Deliverables: Working code, trained model, evaluation metrics on test set.
Exercise 2: Compare Full Fine-Tuning vs LoRA
Task: Fine-tune the same model using both full fine-tuning and LoRA on the same dataset. Compare final performance, training time, and model size.
Learning goals: Understand trade-offs between methods, see LoRA in action.
Deliverables: Comparison table with metrics and time/memory usage.
Exercise 3: Debug a Failing Fine-Tuning
Task: You're given a broken fine-tuning script. Fix the following issues: learning rate too high (NaN loss), data loading bug, incorrect model loading.
Learning goals: Practical debugging skills, common error recognition.
Exercise 4: Build a Custom Evaluation Metric
Task: For your domain (medical, legal, code, etc.), implement a custom evaluation metric beyond generic accuracy. Integrate it into HuggingFace Trainer.
Learning goals: Design metrics for specific use cases, integrate into training pipeline.
Interview Questions on Fine-Tuning
Why fine-tune instead of prompt engineering?
Fine-tuning adapts model internals to your task. Prompts are limited to context window. Fine-tuning gives better performance on specific tasks with custom data.
How do you prevent overfitting during fine-tuning?
Use validation set, early stopping, weight decay, dropout, data augmentation. Monitor train/val loss divergence. Fine-tune fewer epochs.
What's the difference between full fine-tuning and LoRA?
Full: updates all weights. LoRA: updates only low-rank adapters (0.1% of params). LoRA is more efficient but slightly lower quality (95% of full performance).
When would you use QLoRA over LoRA?
When memory is critical. QLoRA quantizes base model to 4-bit, reducing memory by 10x. Trade-off: slower training. Use if you only have consumer GPU.
How do you choose learning rate for fine-tuning?
Typically 1e-5 to 5e-4. Much lower than pre-training (5e-4 to 1e-3). Start with 2e-5, use learning rate finder or grid search if performance is poor.
What's the minimum data needed for fine-tuning?
50-100 examples can improve over base model. 500-2000 gives good results. 5000+ gives very good results. Depends on task complexity.
How do you know if a fine-tuned model is better than base?
Evaluate on task-specific metrics and domain-specific test set. Not on generic benchmarks. Compare inference speed and memory too.
Should you freeze earlier layers during fine-tuning?
Usually no. Unfrozen early layers can adapt to your domain better. Only freeze if you have very little data. Modern practice: fine-tune all layers.
Frequently Asked Questions
How long does fine-tuning take? ▼
Depends on method and hardware. Full fine-tuning on 7B model: 4-24 hours on H100/A100. LoRA: 1-4 hours. QLoRA: 2-8 hours on consumer GPU.
Can I fine-tune a model on CPU? ▼
Yes, but very slow (10-100x slower than GPU). Not practical for large models. Use Google Colab or cloud GPUs for practical fine-tuning.
How much data do I need? ▼
As little as 50 examples can help. But sweet spot is 500-2000 for good results. More data helps up to a point (diminishing returns around 10,000-50,000 examples).
Should I use LoRA or full fine-tuning? ▼
Start with LoRA. It's faster, cheaper, and good enough for most tasks (95% of full FT quality). Use full FT only if LoRA performance is insufficient and you have budget.
How do I know if I'm overfitting? ▼
Validation loss increases while training loss decreases. Model memorizes training examples instead of learning generalizable patterns. Use early stopping or regularization.
Can I fine-tune on multiple GPUs? ▼
Yes. HuggingFace Trainer supports DistributedDataParallel. Just set up your environment and Trainer handles the rest.
What's the difference between fine-tuning and transfer learning? ▼
Transfer learning is the broad concept. Fine-tuning is a specific transfer learning approach. All fine-tuning is transfer learning, but not all transfer learning is fine-tuning.
Can I save multiple adapters for different tasks? ▼
Yes with PEFT. Train different LoRA adapters, save them separately, load the right one at inference based on the task.
How do I serve a fine-tuned model in production? ▼
Merge LoRA weights into base model, then use standard inference (FastAPI, TorchServe, vLLM, etc.). Or quantize for faster inference.
Is fine-tuning still relevant with in-context learning? ▼
Yes. Large context windows help, but fine-tuning still beats few-shot for specific domains. Hybrid: few-shot + retrieval + fine-tuning often wins.
Summary: Key Takeaways
Fine-tuning is transfer learning
Adapt pre-trained models to your task. Requires 10-100x less data than training from scratch.
LoRA is the practical default
Fine-tune with 0.1-1% of parameters. 95% of full FT performance at 1/10 the cost. PEFT library makes it easy.
Data quality matters most
100 perfect examples beat 10,000 noisy ones. Spend time cleaning, deduplicating, validating.
Use validation sets always
Monitor validation metrics during training. Detect overfitting early with early stopping.
Start small and iterate
Begin with 100-200 examples and LoRA. Increase data and switch methods if needed.
Lower learning rates for fine-tuning
Use 1e-5 to 5e-4, not 5e-4 to 1e-3. Pre-trained weights are delicate.
The Fine-Tuning Workflow
1. Prepare Data
Clean, split, validate (100-2000 examples)
โ
2. Choose Method
LoRA (default) or full FT or QLoRA
โ
3. Setup Training
Learning rate, batch size, epochs
โ
4. Monitor & Evaluate
Watch for overfitting, use early stopping
โ
5. Deploy
Save, merge (if LoRA), serve with FastAPI/vLLM
Further Resources
Libraries and Tools
- HuggingFace Transformers: https://github.com/huggingface/transformers โ industry standard library
- PEFT: https://github.com/huggingface/peft โ parameter efficient fine-tuning methods
- TRL: https://github.com/huggingface/trl โ RLHF, DPO, and alignment
- vLLM: https://github.com/lm-sys/vllm โ fast LLM inference
- Optuna: https://optuna.org โ hyperparameter optimization
- Weights & Biases: https://wandb.ai โ experiment tracking
Papers to Read
Practical Guides
- HuggingFace Course: https://huggingface.co/course โ free comprehensive course
- LoRA Fine-Tuning Guide: https://huggingface.co/docs/peft/conceptual_guides/lora
- Llama 2 Fine-Tuning: Meta's official guide and examples
- HuggingFace Blog: Regular posts on fine-tuning techniques and best practices
Communities
- HuggingFace Forums: Active community, questions answered by maintainers
- r/MachineLearning: Reddit community with fine-tuning discussions
- Hugging Face Slack: Real-time chat with researchers and practitioners
- Open LLM Leaderboard: See what fine-tuning techniques lead to SOTA
Next Steps
Now that you understand fine-tuning fundamentals:
- Pick a task and dataset that interests you
- Start with LoRA fine-tuning on a small model (7B-13B)
- Experiment with hyperparameters using the techniques here
- Share results and get feedback from the community
- When comfortable, explore RLHF/DPO for alignment