Sections

Fine-Tuning: Adapting Pre-Trained Models

Fine-tuning is the process of taking a pre-trained neural network model and adapting it to perform well on a specific downstream task. Instead of training a model from scratch (which requires massive amounts of data and compute), you leverage the knowledge already captured by the pre-trained model and only adjust its parameters for your specific domain or task.

Pre-trained models like GPT-3, BERT, Mistral, and Llama have learned rich representations of language from enormous amounts of text. Fine-tuning lets you transfer this knowledge to your specific problem โ€” customer support, medical diagnosis, code generation, or any custom task โ€” with much less data and compute than training from scratch.

Core Concepts

Transfer Learning

Leverage knowledge learned on large datasets for your specific task, reducing data and compute requirements by 10-100x.

Parameter Efficiency

Modern methods like LoRA fine-tune only a fraction of parameters (0.1-1%), making it practical to run on consumer GPUs.

Task Adaptation

Adapt models for instruction-following, classification, generation, Q&A, and more through specialized tuning approaches.

Alignment

Use RLHF and DPO to align models with human preferences and values, reducing hallucinations and harmful outputs.

Key Difference: Pre-training vs Fine-tuning

Pre-training: Training on massive datasets (billions of tokens) to learn general language patterns. Expensive, done rarely by labs like OpenAI/Anthropic.

Fine-tuning: Training on your specific task/domain dataset (thousands to millions of examples). Cheap, can be done by anyone with a GPU.

Why Fine-Tuning Works

Fine-tuning succeeds because:

  • Early layers already know useful features: Early transformer layers have learned to extract phonemes, syntax, named entities, etc. You don't need to relearn these.
  • Task-specific information lives in later layers: Later layers learn task-specific patterns. Fine-tuning primarily adjusts these layers.
  • Smooth loss landscape: Pre-trained models start near a good solution, so fine-tuning rarely gets stuck in bad local minima.
  • Data efficiency: You need maybe 100-1000 examples for fine-tuning vs millions for training from scratch.

Why Fine-Tuning Matters

Fine-tuning is the bridge between cutting-edge research models and practical business applications. It's how companies deploy AI on their specific problems without massive R&D budgets.

The Economics of Fine-Tuning

Training from Scratch (BERT-scale)
$500K+ GPUs, 3-6 months
Fine-tuning Full (GPT-2)
$2-10K, 1-3 days
LoRA (GPT-3.5 scale)
$50-500, 1-4 hours
Prompt Engineering Only
Free, 0-1 hours

Approximate cost/time for adapting a model to your task

Real-World Impact

Customer Support Chatbot

A company fine-tuned Llama-2 on 5,000 internal support tickets and documentation. The fine-tuned model answered customer questions with 92% accuracy, reducing support costs by 40%.

Medical Diagnosis

Researchers fine-tuned PubMedBERT on 10,000 medical notes to predict diagnoses. The fine-tuned model outperformed specialists on rare disease detection by 15%.

Code Generation

A startup fine-tuned CodeLlama on their proprietary codebase and coding standards. The model now generates code that passes tests 78% of the time, vs 35% for the base model.

Methods Comparison: Cost vs Performance

Method GPU Memory Training Time Performance Customization
Prompt Engineering GPU not needed Hours (manual) Low-Medium Limited
Full Fine-Tuning 24-80 GB Hours-Days Very High Maximal
LoRA 8-16 GB 30min-4 hours High Very High
QLoRA 4-8 GB 1-6 hours High Very High
IA3 2-4 GB 30min-2 hours Medium-High High

Key Insight: Full fine-tuning gives maximum performance but is expensive. LoRA gives 95% of full fine-tuning performance at 1/10 the cost. The right method depends on your constraints and performance requirements.

Full Fine-Tuning with HuggingFace Trainer

Full fine-tuning updates all parameters of the model. It's the most flexible approach and gives the best performance, but requires the most compute. Let's build a complete example using the HuggingFace Transformers library.

Setup and Dependencies

Python โ€” Install Dependencies
pip install torch transformers datasets evaluate scikit-learn pip install peft accelerate wandb pip install optuna

Step 1: Load and Prepare Data

Python โ€” Loading and Preprocessing Data
from datasets import load_dataset from transformers import AutoTokenizer # Load a dataset (e.g., AG News for text classification) dataset = load_dataset("ag_news") # Load tokenizer from pre-trained model model_name = "bert-base-uncased" tokenizer = AutoTokenizer.from_pretrained(model_name) def preprocess_function(examples): return tokenizer( examples["text"], truncation=True, max_length=512, padding="max_length" ) # Apply tokenization to all examples tokenized_dataset = dataset.map( preprocess_function, batched=True, num_proc=4 # Use 4 CPU cores ) # Split into train/val train_dataset = tokenized_dataset["train"].shuffle(seed=42) eval_dataset = tokenized_dataset["test"] print(f"Training samples: {len(train_dataset)}") print(f"Eval samples: {len(eval_dataset)}") print(f"Feature keys: {train_dataset.column_names}")

Step 2: Load Pre-Trained Model

Python โ€” Load Model for Fine-Tuning
from transformers import AutoModelForSequenceClassification import torch device = torch.device("cuda" if torch.cuda.is_available() else "cpu") print(f"Using device: {device}") # Load model with number of labels for classification task num_labels = 4 # AG News has 4 categories model = AutoModelForSequenceClassification.from_pretrained( model_name, num_labels=num_labels, problem_type="single_label_classification" ).to(device) print(f"Model parameters: {model.num_parameters():,}") print(f"Trainable parameters: {sum(p.numel() for p in model.parameters() if p.requires_grad):,}")

Step 3: Configure Training Arguments

Python โ€” Training Configuration
from transformers import TrainingArguments training_args = TrainingArguments( output_dir="./results", num_train_epochs=3, per_device_train_batch_size=16, per_device_eval_batch_size=32, gradient_accumulation_steps=2, warmup_steps=500, weight_decay=0.01, learning_rate=2e-5, logging_dir="./logs", logging_steps=100, evaluation_strategy="steps", eval_steps=500, save_strategy="steps", save_steps=500, save_total_limit=2, load_best_model_at_end=True, metric_for_best_model="accuracy", greater_is_better=True, fp16=True, # Mixed precision training gradient_checkpointing=True, # Memory efficient report_to=["tensorboard", "wandb"], )

Step 4: Define Metrics and Trainer

Python โ€” Metrics and Trainer Setup
import numpy as np from datasets import load_metric from transformers import Trainer, TrainerCallback # Load accuracy metric metric = load_metric("accuracy") def compute_metrics(eval_pred): predictions, labels = eval_pred predictions = np.argmax(predictions, axis=1) return metric.compute(predictions=predictions, references=labels) # Create trainer trainer = Trainer( model=model, args=training_args, train_dataset=train_dataset, eval_dataset=eval_dataset, compute_metrics=compute_metrics, callbacks=[], ) print("Trainer initialized and ready for training")

Step 5: Train and Evaluate

Python โ€” Run Training
# Train the model train_result = trainer.train() # Evaluate on test set eval_results = trainer.evaluate(eval_dataset) print(f"\nEvaluation Results: {eval_results}") # Save the fine-tuned model trainer.save_model("./fine_tuned_model") print("Model saved to ./fine_tuned_model")

Step 6: Make Predictions

Python โ€” Inference with Fine-Tuned Model
from transformers import pipeline # Load the fine-tuned model pipe = pipeline( "text-classification", model="./fine_tuned_model", device=device ) # Make predictions texts = [ "Apple releases new iPhone with better battery", "Stock market crashes amid economic concerns", "Sports: Team wins championship after overtime" ] predictions = pipe(texts, batch_size=8) for text, pred in zip(texts, predictions): print(f"Text: {text}") print(f" -> Label: {pred['label']}, Score: {pred['score']:.4f}\n")

Key Parameters Explained

Learning Rate (2e-5)

Fine-tuning uses much smaller learning rates than training from scratch (typically 1e-5 to 5e-5). The model already knows a lot โ€” we're making small adjustments.

Warmup Steps

Gradually increase learning rate at the start. Without warmup, large gradient updates can throw the model off the pre-trained solution.

Weight Decay

L2 regularization penalty. Prevents overfitting when fine-tuning on small datasets.

Gradient Accumulation

Accumulate gradients over multiple batches before updating. Simulates larger batch size without needing more GPU memory.

Memory and Speed Optimizations

  • Gradient Checkpointing: Trade compute for memory. Store fewer activations during forward pass, recompute during backward. Reduces memory by ~50%, increases training time by ~20-30%.
  • Mixed Precision (fp16): Use 16-bit floats instead of 32-bit. Faster and less memory, nearly identical results.
  • Gradient Accumulation: Simulate larger batch sizes without more GPU memory.
  • LoRA (next section): Update only a tiny fraction of parameters.

LoRA: Efficient Fine-Tuning with Low-Rank Adaptation

LoRA (Low-Rank Adaptation) is a technique that reduces the number of trainable parameters by 100-1000x while maintaining most of the performance of full fine-tuning. Instead of updating all parameters, LoRA adds small, trainable "adapter" matrices to the model.

The LoRA Idea

The key insight: adaptation to a downstream task doesn't require updating all parameters. Instead, we can express the weight updates as a low-rank decomposition.

LoRA Weight Update
For a weight matrix W (dimensions d ร— k), instead of: W' = W + ฮ”W (d ร— k update, expensive) We compute: W' = W + AB^T where A is (d ร— r) and B is (k ร— r) with r << min(d, k) (the rank) This reduces parameters from dยทk to rยท(d+k) Example: For W = 4096 ร— 4096: Full update: 16.8M parameters LoRA (r=8): 65.5K parameters (0.4%!)

LoRA Setup with PEFT

Python โ€” LoRA Configuration
from peft import LoraConfig, get_peft_model # Configure LoRA lora_config = LoraConfig( r=8, # LoRA rank lora_alpha=16, # Scaling factor (usually 2 * r) target_modules=["q_proj", "v_proj"], # Which modules to apply LoRA to lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" # Or "SEQ_2_SEQ_LM", "QUESTION_ANS", etc. ) # Wrap model with LoRA model = get_peft_model(model, lora_config) model.print_trainable_parameters() # Output: # trainable params: 4194304 || all params: 124439552 || trainable%: 3.37%

Full LoRA Fine-Tuning Example

Python โ€” Complete LoRA Fine-Tuning
from transformers import Trainer, TrainingArguments from peft import LoraConfig, get_peft_model # 1. Load base model and tokenizer model = AutoModelForSequenceClassification.from_pretrained( "bert-base-uncased", num_labels=4 ) tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased") # 2. Apply LoRA lora_config = LoraConfig( r=8, lora_alpha=16, target_modules=["query", "value"], lora_dropout=0.05, bias="none", task_type="SEQ_CLS" ) model = get_peft_model(model, lora_config) # 3. Setup training (same as full fine-tuning) training_args = TrainingArguments( output_dir="./lora_results", num_train_epochs=3, per_device_train_batch_size=32, # Can use larger batch learning_rate=1e-4, # Slightly higher than full FT fp16=True, logging_steps=100, evaluation_strategy="steps", eval_steps=500, ) # 4. Create and run trainer trainer = Trainer( model=model, args=training_args, train_dataset=train_dataset, eval_dataset=eval_dataset, compute_metrics=compute_metrics, ) trainer.train() # 5. Save LoRA weights (only ~4MB!) model.save_pretrained("./lora_weights") print("LoRA weights saved")

LoRA Hyperparameters

Parameter Typical Range Effect
r (rank) 4, 8, 16, 32 Higher = more capacity but slower. 8 is often optimal.
lora_alpha 8, 16, 32 Scaling factor. Usually 2x rank. Controls adapter strength.
lora_dropout 0.05, 0.1, 0.15 Regularization. Prevents overfitting on small datasets.
target_modules ["q_proj", "v_proj"] Which layers to apply LoRA. Keys have less impact than query/value.

QLoRA: Even More Efficient

QLoRA combines LoRA with quantization to reduce memory even further. It quantizes the base model to 4-bit, then adds tiny LoRA adapters on top.

Python โ€” QLoRA Configuration
from transformers import BitsAndBytesConfig from peft import LoraConfig, get_peft_model # Quantization config (4-bit) bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_use_double_quant=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) # Load quantized model model = AutoModelForCausalLM.from_pretrained( "meta-llama/Llama-2-7b-hf", quantization_config=bnb_config, device_map="auto", ) # Apply LoRA on top lora_config = LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "v_proj", "k_proj", "o_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" ) model = get_peft_model(model, lora_config) # Now training uses only ~8GB VRAM for 7B model! print(model.print_trainable_parameters())

LoRA Insight: Most of the fine-tuning updates are low-rank! This is why LoRA works so well. The model is doing something fundamentally different for your task, but that difference lives in a low-rank subspace.

Loading and Using LoRA Models

Python โ€” Inference with LoRA
from peft import AutoPeftModelForSequenceClassification # Load base model + LoRA weights automatically model = AutoPeftModelForSequenceClassification.from_pretrained( "path/to/lora_weights" ) # Option 1: Use as-is for inference outputs = model(input_ids, attention_mask=attention_mask) # Option 2: Merge LoRA into base weights (one-time cost) merged_model = model.merge_and_unload() # Save merged model (can use with any inference framework) merged_model.save_pretrained("./merged_model")

Instruction Tuning: Teaching Models to Follow Instructions

Instruction tuning teaches language models to follow user instructions and respond appropriately. Instead of just predicting the next token, the model learns to understand tasks described in natural language and execute them correctly.

Why Instruction Tuning?

Pre-trained language models are trained to predict the next token. They don't inherently understand task descriptions or follow instructions. A model might complete "The capital of France is" with accurate information, but fail at "Answer in French: What is the capital of France?"

Instruction tuning aligns the model with human expectations: given a task description (prompt), generate the correct output.

Dataset Format

Python โ€” Instruction Tuning Dataset Format
{ "instruction": "Classify the sentiment of this review.", "input": "This movie was absolutely terrible. Boring, slow, and a waste of time.", "output": "negative" } { "instruction": "Summarize the following article in 2-3 sentences.", "input": "New research shows that regular exercise...", "output": "The study found that..." } { "instruction": "Translate to French", "input": "Hello, how are you?", "output": "Bonjour, comment allez-vous?" } { "instruction": "Write a Python function that...", "input": "that takes a list and returns unique elements", "output": "def unique(lst):\n return list(set(lst))" }

Building an Instruction Dataset

Python โ€” Creating Instruction Tuning Data
import json from datasets import Dataset # Create instruction tuning examples examples = [ { "instruction": "Classify the sentiment", "input": "This product is amazing!", "output": "positive" }, { "instruction": "Classify the sentiment", "input": "Worst purchase ever", "output": "negative" }, # ... more examples ] # Format for training def format_instruction(example): return { "text": f"""Instruction: {example['instruction']} Input: {example['input']} Output: {example['output']}<|endoftext|>""" } formatted = [format_instruction(ex) for ex in examples] dataset = Dataset.from_dict({ "text": [ex["text"] for ex in formatted] }) print(f"Dataset size: {len(dataset)} examples")

Fine-Tuning for Instruction Following

Python โ€” Instruction Tuning with Trainer
from transformers import ( AutoModelForCausalLM, AutoTokenizer, Trainer, TrainingArguments, ) # Load model model_name = "gpt2" # Or mistral, llama, etc. model = AutoModelForCausalLM.from_pretrained(model_name) tokenizer = AutoTokenizer.from_pretrained(model_name) # Prepare dataset def preprocess_function(examples): return tokenizer( examples["text"], truncation=True, max_length=512, ) train_dataset = dataset.map(preprocess_function, batched=True) # Training arguments (instruction tuning is lighter than pretraining) args = TrainingArguments( output_dir="./instruction_model", num_train_epochs=3, per_device_train_batch_size=8, learning_rate=5e-5, warmup_steps=100, weight_decay=0.01, logging_steps=10, save_strategy="epoch", fp16=True, ) # Train trainer = Trainer( model=model, args=args, train_dataset=train_dataset, ) trainer.train() model.save_pretrained("./instruction_tuned_model")

Inference: Using the Instruction-Tuned Model

Python โ€” Generating Outputs with Instructions
from transformers import pipeline pipe = pipeline( "text-generation", model="./instruction_tuned_model", device=0 ) # Create prompt in same format as training prompt = """Instruction: Classify sentiment Input: This is great! Output:""" output = pipe( prompt, max_length=50, num_return_sequences=1, temperature=0.7, top_p=0.9, ) print(output[0]["generated_text"])

Best Practices for Instruction Data

Diversity

Include variety: different tasks, domains, difficulty levels. A dataset with only classification gets overtrained on that task.

Quality > Quantity

100 high-quality instruction examples beat 10,000 low-quality ones. Spend time curating.

Clear Instructions

Write instructions that a human could understand. Avoid ambiguity. Bad: 'Do something'. Good: 'Classify the sentiment as positive, negative, or neutral.'

Balanced Outputs

If possible, balance output lengths and task types. Prevents bias toward one type of response.

Instruction Tuning vs Full Fine-Tuning

Aspect Instruction Tuning Full Fine-Tuning
Task Type General purpose (any task) Specific task
Data Format instruction + input โ†’ output input โ†’ output
Model Flexibility High (understands new tasks) Low (locked to task)
Data Required 500-5000 diverse examples 100-1000 task-specific examples
Best For Building helpful assistants Specific applications

Alignment: RLHF and Direct Preference Optimization (DPO)

After instruction tuning, models often hallucinate, refuse valid requests, or produce harmful content. RLHF (Reinforcement Learning from Human Feedback) and DPO (Direct Preference Optimization) align models with human preferences and values.

The Problem: Why Fine-Tuning Isn't Enough

Consider these issues:

  • Hallucination: Model confidently states false information.
  • Refusal: Model refuses to help with legitimate requests.
  • Helpfulness: Model gives technically correct but unhelpful answers.
  • Harmfulness: Model generates toxic, biased, or dangerous content.

Standard fine-tuning on examples doesn't address these โ€” we need to optimize for preferences.

RLHF: Reinforcement Learning from Human Feedback

RLHF has 4 phases:

1. Supervised Fine-Tuning

Fine-tune on high-quality instruction examples (previous section).

2. Reward Modeling

Collect human preferences (good vs bad outputs). Train a model to predict which output humans prefer.

3. Reinforcement Learning

Use the reward model to train the language model. Maximize expected reward while staying close to pre-trained model.

4. Red Teaming

Find and fix failure modes. Repeat until model is safe and helpful.

Reward Model Training

Python โ€” Training a Reward Model
from transformers import AutoModelForSequenceClassification, Trainer import torch # Dataset format: prompt + chosen response + rejected response # Collect via human annotations preference_data = [ { "prompt": "Tell me a joke", "chosen": "Why did the chicken cross the road? To get to the other side!", "rejected": "I cannot tell jokes." }, # More examples... ] class PairwiseDataset(Dataset): def __init__(self, data, tokenizer): self.tokenizer = tokenizer self.data = data def __len__(self): return len(self.data) def __getitem__(self, idx): example = self.data[idx] # Tokenize chosen response chosen = self.tokenizer( example["prompt"] + example["chosen"], max_length=512, truncation=True, return_tensors="pt" ) # Tokenize rejected response rejected = self.tokenizer( example["prompt"] + example["rejected"], max_length=512, truncation=True, return_tensors="pt" ) return { "chosen_input_ids": chosen["input_ids"].squeeze(), "chosen_attention_mask": chosen["attention_mask"].squeeze(), "rejected_input_ids": rejected["input_ids"].squeeze(), "rejected_attention_mask": rejected["attention_mask"].squeeze(), } # Train reward model model = AutoModelForSequenceClassification.from_pretrained( "gpt2", num_labels=1 # Single scalar reward ) # Custom loss: maximize difference between chosen and rejected def reward_loss(model_output_chosen, model_output_rejected): return -torch.nn.functional.logsigmoid( model_output_chosen.logits - model_output_rejected.logits ).mean() print("Reward model training setup ready")

Direct Preference Optimization (DPO)

DPO is a simpler alternative to RLHF that doesn't require training a separate reward model. Instead, it directly optimizes the language model using preference data.

DPO Objective
Instead of: 1. Training reward model 2. RL to maximize reward DPO directly minimizes: L_DPO = -log ฯƒ(ฮฒ log(ฯ€_ฮธ(y|x) / ฯ€_ref(y|x)) - ฮฒ log(ฯ€_ฮธ(y'|x) / ฯ€_ref(y'|x))) Where: ฯƒ = sigmoid function ฮฒ = temperature controlling preference strength ฯ€_ฮธ = our model ฯ€_ref = reference (pre-trained) model y = chosen response y' = rejected response

DPO Implementation

Python โ€” DPO Fine-Tuning with TRL
from trl import DPOTrainer, DPOConfig from transformers import AutoModelForCausalLM, AutoTokenizer # Prepare preference data preference_dataset = [ { "prompt": "Tell me a joke", "chosen": "Why did the chicken cross the road?", "rejected": "I can't tell jokes.", }, # More examples... ] # Load models model = AutoModelForCausalLM.from_pretrained("gpt2") tokenizer = AutoTokenizer.from_pretrained("gpt2") # DPO training dpo_config = DPOConfig( beta=0.1, # Temperature for preference learning_rate=1e-5, num_train_epochs=3, per_device_train_batch_size=4, output_dir="./dpo_model", ) trainer = DPOTrainer( model=model, args=dpo_config, train_dataset=preference_dataset, tokenizer=tokenizer, ) trainer.train() model.save_pretrained("./dpo_model")

RLHF vs DPO Comparison

Aspect RLHF DPO
Complexity High (4 phases) Low (direct optimization)
Training Cost Expensive (RL + reward model) Cheap (direct training)
Reward Model Separate model needed No separate model
Stability Can be unstable More stable
Performance Excellent Excellent
When to Use Production systems needing fine control Most practitioners

Key Insight: DPO has become increasingly popular because it's simpler than RLHF while achieving similar or better results. Unless you have specific reasons for RLHF, start with DPO.

Dataset Preparation for Fine-Tuning

Quality data is everything in fine-tuning. A small dataset of high-quality examples beats a large dataset of noisy examples. Let's cover practical dataset preparation techniques.

Data Sources and Collection

Internal Data

Customer tickets, documentation, past conversations. Most valuable! Domain-specific and private.

Public Datasets

HuggingFace Hub, Kaggle, academic datasets. Free but may not match your domain exactly.

Synthetic Data

Generate examples using GPT-4 or other strong models. Useful for augmentation or when real data is scarce.

Crowdsourcing

Hire annotators to create or label data. Expensive but high quality if managed well.

Data Cleaning and Preprocessing

Python โ€” Cleaning and Validation
import json import re from datasets import Dataset def clean_text(text): '''Remove noise from text''' # Remove HTML tags text = re.sub(r'<[^>]+>', '', text) # Remove extra whitespace text = re.sub(r'\s+', ' ', text).strip() # Remove special characters (but keep punctuation) text = re.sub(r'[^a-zA-Z0-9\s.!?,-]', '', text) return text def validate_example(example): '''Check if example is valid for fine-tuning''' if len(example.get("input", "")) < 5: return False if len(example.get("output", "")) < 5: return False if "input" not in example or "output" not in example: return False return True # Load and clean data with open("raw_data.jsonl") as f: raw_examples = [json.loads(line) for line in f] # Filter and clean cleaned = [] for ex in raw_examples: if validate_example(ex): ex["input"] = clean_text(ex["input"]) ex["output"] = clean_text(ex["output"]) cleaned.append(ex) print(f"Kept {len(cleaned)}/{len(raw_examples)} examples") dataset = Dataset.from_dict({ "input": [ex["input"] for ex in cleaned], "output": [ex["output"] for ex in cleaned], })

Train/Validation/Test Split

Python โ€” Creating Dataset Splits
from sklearn.model_selection import train_test_split # Split: 80% train, 10% val, 10% test train, temp = train_test_split(cleaned, test_size=0.2, random_state=42) val, test = train_test_split(temp, test_size=0.5, random_state=42) print(f"Train: {len(train)} | Val: {len(val)} | Test: {len(test)}") # Create datasets train_dataset = Dataset.from_dict({ "input": [ex["input"] for ex in train], "output": [ex["output"] for ex in train], }) val_dataset = Dataset.from_dict({ "input": [ex["input"] for ex in val], "output": [ex["output"] for ex in val], }) test_dataset = Dataset.from_dict({ "input": [ex["input"] for ex in test], "output": [ex["output"] for ex in test], }) # Save splits train_dataset.save_to_disk("./train_data") val_dataset.save_to_disk("./val_data") test_dataset.save_to_disk("./test_data")

Data Augmentation Techniques

Python โ€” Augmenting Training Data
import random def paraphrase_instruction(instruction): '''Simple paraphrasing of instructions''' templates = [ "Can you {}?", "Please {}.", "I need you to {}.", "Would you {}?", "How would you {}?", ] # Extract key verbs and use different template base = instruction.lower() template = random.choice(templates) return template.format(base) def add_variations(dataset): '''Augment dataset with variations''' augmented = [] for example in dataset: # Original augmented.append(example) # Paraphrased instruction (if it's an instruction task) if "instruction" in example: aug_ex = example.copy() aug_ex["instruction"] = paraphrase_instruction(example["instruction"]) augmented.append(aug_ex) return augmented # Augment training data augmented_train = add_variations(train) print(f"Training size after augmentation: {len(augmented_train)}")

Data Analysis and Debugging

Python โ€” Analyzing Your Dataset
import numpy as np from collections import Counter # Length statistics input_lengths = [len(ex["input"].split()) for ex in dataset] output_lengths = [len(ex["output"].split()) for ex in dataset] print(f"Input length - Mean: {np.mean(input_lengths):.0f}, " f"Min: {np.min(input_lengths)}, Max: {np.max(input_lengths)}") print(f"Output length - Mean: {np.mean(output_lengths):.0f}, " f"Min: {np.min(output_lengths)}, Max: {np.max(output_lengths)}") # Check for duplicates inputs = [ex["input"] for ex in dataset] print(f"Duplicates: {len(inputs) - len(set(inputs))}") # Class balance (for classification) if "label" in dataset[0]: labels = [ex["label"] for ex in dataset] print("Label distribution:", Counter(labels))

Dataset Size Guidelines

Fine-Tuning Method Minimum Data Recommended Optimal
Full Fine-Tuning (Classification) 100 500-1000 5000+
Full Fine-Tuning (Generation) 200 1000-2000 10000+
LoRA (Classification) 50 200-500 2000+
LoRA (Generation) 100 500-1000 5000+
Instruction Tuning 200 1000-2000 5000+

Rule of Thumb: You need at least 100 examples to improve over a pre-trained model. 500-2000 gives good results for most tasks. More data always helps, but with diminishing returns.

Evaluation During Training: Metrics and Monitoring

You can't improve what you don't measure. Proper evaluation during fine-tuning is crucial for understanding progress and preventing overfitting.

Classification Metrics

Python โ€” Computing Classification Metrics
from sklearn.metrics import ( accuracy_score, precision_recall_fscore_support, confusion_matrix, roc_auc_score, classification_report, ) import numpy as np def compute_classification_metrics(predictions, labels, num_classes=None): '''Compute standard classification metrics''' # Accuracy accuracy = accuracy_score(labels, predictions) # Precision, Recall, F1 precision, recall, f1, _ = precision_recall_fscore_support( labels, predictions, average="weighted" ) # Per-class metrics report = classification_report(labels, predictions, output_dict=True) # Confusion matrix cm = confusion_matrix(labels, predictions) return { "accuracy": accuracy, "precision": precision, "recall": recall, "f1": f1, "per_class": report, "confusion_matrix": cm, } # Usage in training predictions = np.argmax(logits, axis=1) metrics = compute_classification_metrics(predictions, labels) print(f"Accuracy: {metrics['accuracy']:.4f}") print(f"F1: {metrics['f1']:.4f}")

Generation Metrics

Python โ€” Evaluating Text Generation
from rouge_score import rouge_scorer from nltk.translate.bleu_score import corpus_bleu import nltk # Download NLTK data nltk.download('punkt') def compute_generation_metrics(predictions, references): '''Compute metrics for generation tasks (summarization, translation)''' # ROUGE (for summarization) rouge = rouge_scorer.RougeScorer(['rouge1', 'rougeL'], use_stemmer=True) rouge_scores = [rouge.score(ref, pred) for pred, ref in zip(predictions, references)] avg_rouge1 = np.mean([s['rouge1'].fmeasure for s in rouge_scores]) # BLEU (for translation) references_tokens = [[ref.split()] for ref in references] predictions_tokens = [pred.split() for pred in predictions] bleu = corpus_bleu(references_tokens, predictions_tokens) # Exact Match (for QA, simple tasks) exact_match = sum([p.strip() == r.strip() for p, r in zip(predictions, references)]) / len(predictions) return { "rouge1": avg_rouge1, "bleu": bleu, "exact_match": exact_match, } # Usage preds = ["The cat sat on the mat", "Dogs are animals"] refs = ["A cat was sitting on the mat", "Dogs are mammals"] metrics = compute_generation_metrics(preds, refs) print(f"ROUGE-1: {metrics['rouge1']:.4f}")

Custom Evaluation Metrics

Python โ€” Creating Domain-Specific Metrics
def evaluate_code_generation(predictions, references): '''Custom metric: code generation accuracy''' scores = [] for pred, ref in zip(predictions, references): try: # Try to execute both and compare output pred_result = eval(pred) ref_result = eval(ref) scores.append(1.0 if pred_result == ref_result else 0.0) except: # Execution failed scores.append(0.0) return np.mean(scores) def evaluate_medical_diagnosis(predictions, references, task="sensitivity"): '''Custom metric: sensitivity/specificity trade-off''' from sklearn.metrics import recall_score, precision_score if task == "sensitivity": # Maximize recall (catch all positives) return recall_score(references, predictions) elif task == "specificity": # Maximize specificity (avoid false alarms) return recall_score(references, predictions, pos_label=0) # Use in evaluation def compute_metrics(eval_pred): predictions, labels = eval_pred predictions = np.argmax(predictions, axis=1) # Standard metrics accuracy = accuracy_score(labels, predictions) # Custom metrics sensitivity = evaluate_medical_diagnosis(predictions, labels, "sensitivity") return { "accuracy": accuracy, "sensitivity": sensitivity, }

Monitoring Training with Logging

Python โ€” Setup Logging and Monitoring
import wandb from transformers import TrainingArguments, Trainer # Initialize wandb wandb.init(project="fine-tuning", name="my-experiment") # Configure logging training_args = TrainingArguments( output_dir="./results", num_train_epochs=3, per_device_train_batch_size=16, logging_dir="./logs", logging_steps=100, evaluation_strategy="steps", eval_steps=500, # Wandb integration report_to=["wandb"], run_name="gpt2-finetuning-v1", ) # Create trainer with callbacks trainer = Trainer( model=model, args=training_args, train_dataset=train_dataset, eval_dataset=eval_dataset, compute_metrics=compute_metrics, callbacks=[], ) # Log custom metrics wandb.log({ "dataset_size": len(train_dataset), "model_parameters": model.num_parameters(), }) trainer.train()

Detecting Overfitting

Look for these signs during training:

  • Training loss keeps decreasing: Good, model is learning.
  • Validation loss increases: Overfitting! Regularize more (higher dropout, weight decay, data augmentation).
  • Training and validation loss diverge: Model overfitting. Too many epochs or not enough data.
  • Metrics plateau: Diminishing returns. Training longer won't help. Try more data, adjust hyperparameters, or accept current performance.

Early Stopping

Stop training when validation loss stops improving for N evaluations. Prevents overfitting and saves compute. Most trainers support this via callbacks.

Hyperparameter Search and Optimization

Fine-tuning success depends heavily on hyperparameters: learning rate, batch size, warmup steps, weight decay, etc. Random hyperparameter search is inefficient. Optuna enables systematic optimization.

Common Hyperparameters to Tune

Hyperparameter Typical Range Impact Tuning Order
Learning Rate 1e-5 to 5e-4 Very High 1st
Batch Size 8 to 128 High 2nd
Warmup Steps 0 to 1000 Medium 3rd
Weight Decay 0 to 0.1 Medium 3rd
Epochs 2 to 10 Medium 4th

Optuna-Based Hyperparameter Search

Python โ€” Systematic Hyperparameter Tuning
import optuna from transformers import Trainer, TrainingArguments import numpy as np def objective(trial): '''Objective function for Optuna to optimize''' # Suggest hyperparameters learning_rate = trial.suggest_float("learning_rate", 1e-5, 5e-4, log=True) batch_size = trial.suggest_int("batch_size", 8, 32) warmup_steps = trial.suggest_int("warmup_steps", 0, 1000, step=100) weight_decay = trial.suggest_float("weight_decay", 0.0, 0.1) # Create trainer with these hyperparameters training_args = TrainingArguments( output_dir=f"./trial_{trial.number}", num_train_epochs=3, per_device_train_batch_size=batch_size, per_device_eval_batch_size=32, learning_rate=learning_rate, warmup_steps=warmup_steps, weight_decay=weight_decay, evaluation_strategy="steps", eval_steps=500, save_strategy="no", # Don't save to speed up search logging_steps=100, fp16=True, ) trainer = Trainer( model=model, args=training_args, train_dataset=train_dataset, eval_dataset=eval_dataset, compute_metrics=compute_metrics, ) # Train and get best validation metric trainer.train() eval_result = trainer.evaluate() return eval_result["eval_accuracy"] # Create and run study study = optuna.create_study(direction="maximize") study.optimize(objective, n_trials=20) # Try 20 different configurations # Get best hyperparameters best_trial = study.best_trial print(f"Best accuracy: {best_trial.value:.4f}") print(f"Best hyperparameters: {best_trial.params}")

Grid Search vs Random Search vs Bayesian Optimization

Method Efficiency Complexity Best For
Grid Search Low Low Small search spaces
Random Search Medium Low Quick experiments
Bayesian Opt (Optuna) High Medium Production tuning

Learning Rate Scheduling

Python โ€” Advanced Learning Rate Scheduling
from transformers import TrainingArguments import math # Different scheduling strategies training_args = TrainingArguments( output_dir="./results", # Learning rate schedule options: # 1. Constant with warmup (default) learning_rate=2e-5, warmup_steps=500, # 2. Linear decay (specify via lr_scheduler_type) lr_scheduler_type="linear", # 3. Cosine decay lr_scheduler_type="cosine", # 4. Cosine with hard restarts (learning rate resets during training) lr_scheduler_type="cosine_with_restarts", # Number of hard restarts num_cycles=3, ) # Custom schedule via callback class CustomScheduleCallback: def on_step_end(self, args, state, control, **kwargs): # Custom logic to adjust learning rate pass

Learning Rate Finder

Python โ€” Finding Optimal Learning Rate
import matplotlib.pyplot as plt import numpy as np def learning_rate_finder(model, train_loader, device, start_lr=1e-6, end_lr=1e-1, num_iterations=100): '''Find optimal learning rate by training with increasing LR''' lrs = np.logspace(np.log10(start_lr), np.log10(end_lr), num_iterations) losses = [] for lr in lrs: # Set learning rate for param_group in optimizer.param_groups: param_group['lr'] = lr # Train for one batch model.train() batch = next(iter(train_loader)) outputs = model(batch) loss = outputs.loss optimizer.zero_grad() loss.backward() optimizer.step() losses.append(loss.item()) # Plot results plt.figure(figsize=(10, 6)) plt.xscale('log') plt.plot(lrs, losses) plt.xlabel('Learning Rate') plt.ylabel('Loss') plt.title('Learning Rate Finder') plt.savefig('./lr_finder.png') # Find steepest descent best_lr = lrs[np.argmin(losses)] print(f"Suggested learning rate: {best_lr:.2e}") return best_lr

Comprehensive Comparison: Full vs LoRA vs QLoRA vs Adapters

Multiple efficient fine-tuning methods exist. Each has trade-offs in memory, speed, and performance. This section compares them across practical dimensions.

Visual Comparison: Memory vs Performance

Full Fine-Tuning
100% Performance, 100% Memory
LoRA (r=8)
95% Performance, 30% Memory
QLoRA
92% Performance, 10% Memory
IA3
85% Performance, 5% Memory
Prompt Tuning
70% Performance, 2% Memory

Approximate relative performance and memory usage for a 7B model

Detailed Comparison Table

Method Trainable Params GPU Memory Training Time Final Performance Complexity
Full Fine-Tuning 100% 80 GB (7B model) 8-24h 100% Low
LoRA 0.1-1% 24-40 GB 2-8h 95-98% Low
QLoRA 0.1-1% 8-16 GB 4-12h 92-96% Medium
IA3 0.01-0.1% 4-8 GB 2-6h 80-90% Low
Prefix Tuning 0.1-1% 16-24 GB 3-6h 90-95% Medium
Prompt Tuning 0.01-0.1% 2-4 GB 30min-2h 70-85% Low

Which Method to Use?

Full Fine-Tuning

Use when: Maximum performance is critical, you have compute budget, task is very different from pre-training. Gives best results.

LoRA

Use when: Good performance with moderate efficiency needed. Most practitioners should start here. Great balance of quality and cost.

QLoRA

Use when: Limited GPU memory (RTX 3090, RTX 4090). Adds 20-30% training time but saves 5x memory vs LoRA.

IA3

Use when: Extreme memory constraints or many small adaptations needed. Performance trade-off is acceptable.

Combining Methods: Multi-Adapter Fine-Tuning

Advanced: Train multiple task-specific adapters and switch between them at inference time.

Python โ€” Using Multiple Task Adapters
from peft import PeftModel, LoraConfig, get_peft_model # Load base model once base_model = AutoModelForCausalLM.from_pretrained("mistral-7b") # Create adapters for different tasks adapters = {} tasks = ["sentiment", "translation", "summarization"] for task in tasks: lora_config = LoraConfig( r=8, lora_alpha=16, target_modules=["q_proj", "v_proj"], task_type="CAUSAL_LM" ) model = get_peft_model(base_model, lora_config) # Fine-tune for this task... trainer.train() # Save adapter model.save_pretrained(f"./adapters/{task}") adapters[task] = model # At inference, load the right adapter for the task loaded_model = PeftModel.from_pretrained( base_model, f"./adapters/sentiment" ) # Switch adapters without reloading model loaded_model.set_adapter("translation") # Switch to translation adapter

Practical Tip: If you're unsure, start with LoRA. It has great performance, is memory efficient, and is well-supported across frameworks.

Fine-Tuning Best Practices and Common Pitfalls

10 Essential Best Practices

1. Start Small and Iterate

Don't fine-tune on 10,000 examples immediately. Start with 100-200 examples. If performance is poor, the issue is likely not data size but hyperparameters or data quality.

2. Always Use a Validation Set

Monitor validation metrics during training. Without validation, you can't detect overfitting. Use early stopping to prevent wasted compute.

3. Lower Learning Rate Than Pre-training

Fine-tuning uses 10-100x smaller learning rates than pre-training. Typically 1e-5 to 5e-4. Too high and you destroy pre-trained knowledge. Too low and training is painfully slow.

4. Use Mixed Precision Training

Enable fp16=True in training args. Speeds up training by 20-30%, uses half the memory, and accuracy is nearly identical. It's a free win.

5. Data Quality Beats Quantity

100 perfect examples beat 10,000 noisy ones. Spend time cleaning, deduplicating, and validating data. This is often the highest ROI task.

6. Try Different Model Sizes

Bigger models are often better, but smaller models can work for simple tasks. A 7B model fine-tuned on your data might beat a 70B base model with prompts.

7. Monitor for Mode Collapse

If your fine-tuned model generates identical outputs regardless of input, it's collapsed. Increase data diversity, reduce regularization, or lower learning rate.

8. Save Checkpoints Regularly

Don't save only the final model. Save intermediate checkpoints. Sometimes earlier checkpoints are better due to overfitting.

9. Use Gradient Accumulation for Large Batches

If you can't fit large batches in memory, use gradient accumulation. Simulates larger batches with equivalent results but lower memory.

10. Test on Domain-Specific Data

Evaluate on examples from your actual use case, not just public benchmarks. A model that scores 95% on GLUE might perform poorly on your specific task.

Common Pitfalls and Solutions

Fine-tuned model performs worse than base model โ–ผ

Likely causes:

  • Learning rate too high โ€” reduces to 1e-5
  • Not enough training data โ€” gather more examples
  • Data distribution mismatch โ€” ensure data is from your domain
  • Task is too different from pre-training โ€” try different architecture
Training is very slow or OOM errors โ–ผ
  • Reduce batch size (try 2-4 instead of 32)
  • Enable gradient checkpointing: gradient_checkpointing=True
  • Use LoRA instead of full fine-tuning
  • Use QLoRA for even more memory savings
  • Reduce max_seq_length (e.g., 256 instead of 512)
Model overfits on training data โ–ผ
  • Increase weight decay (try 0.01 or 0.1)
  • Increase dropout (lora_dropout=0.1 or higher)
  • Reduce number of epochs (try 1-2 instead of 3-5)
  • Add data augmentation
  • Use early stopping based on validation loss

Advanced Fine-Tuning Techniques

Multi-task Fine-Tuning

Train on multiple tasks simultaneously. The model learns shared representations and often performs better on individual tasks.

Python โ€” Multi-Task Fine-Tuning
from torch.utils.data import ConcatDataset, DataLoader # Combine multiple task datasets sentiment_data = load_dataset("amazon_reviews_multi")["train"] topic_data = load_dataset("ag_news")["train"] # Create multi-task examples def add_task_marker(example, task_name): example["task"] = task_name example["text"] = f"Task: {task_name}\n{example['text']}" return example sentiment_data = sentiment_data.map( lambda x: add_task_marker(x, "sentiment") ) topic_data = topic_data.map( lambda x: add_task_marker(x, "topic") ) # Combine datasets combined = ConcatDataset([sentiment_data, topic_data]) # The model learns to solve both tasks trainer = Trainer(...) trainer.train()

Continued Pre-training (Domain Adaptation)

Sometimes better than fine-tuning: continue pre-training on unlabeled data from your domain before fine-tuning on labeled data.

Python โ€” Domain Adaptation Pipeline
# Phase 1: Continued pre-training on domain data # Use causal language modeling loss on unlabeled data from your domain training_args = TrainingArguments( output_dir="./domain_adapted_model", num_train_epochs=1, per_device_train_batch_size=16, learning_rate=5e-5, fp16=True, ) # Use a dataset of unlabeled text from your domain unlabeled_data = load_dataset("text", data_files="domain_corpus.txt") trainer = Trainer( model=model, args=training_args, train_dataset=unlabeled_data, data_collator=DataCollatorForLanguageModeling(tokenizer, mlm=False), ) trainer.train() # Phase 2: Fine-tune on labeled task data # Now fine-tune the domain-adapted model on your task model = AutoModelForSequenceClassification.from_pretrained( "./domain_adapted_model", num_labels=4 ) trainer = Trainer(...) trainer.train()

Few-Shot In-Context Learning

Alternative to fine-tuning: include examples in the prompt.

Python โ€” Few-Shot Prompting
from transformers import pipeline pipe = pipeline("text-classification", model="gpt2") # Few-shot examples in the prompt prompt = """Classify the sentiment. Example 1: Input: Great movie! Output: positive Example 2: Input: Terrible experience Output: negative Now classify: Input: This is okay. Output:""" output = pipe(prompt, max_length=200) print(output)

Knowledge Distillation

Train a smaller, faster model to mimic a larger fine-tuned model.

Python โ€” Knowledge Distillation Setup
import torch import torch.nn.functional as F def distillation_loss(student_logits, teacher_logits, labels, temperature=3.0, alpha=0.7): '''Combined supervised + distillation loss''' # Supervised loss ce_loss = F.cross_entropy(student_logits, labels) # Distillation loss: KL divergence of softened distributions kl_loss = F.kl_div( F.log_softmax(student_logits / temperature, dim=1), F.softmax(teacher_logits / temperature, dim=1), reduction='batchmean' ) * (temperature ** 2) # Combined loss = alpha * ce_loss + (1 - alpha) * kl_loss return loss # Train student model with teacher guidance for batch in train_loader: student_logits = student_model(batch) with torch.no_grad(): teacher_logits = teacher_model(batch) loss = distillation_loss(student_logits, teacher_logits, batch['labels']) loss.backward() optimizer.step()

Deployment: Saving, Merging, and Serving Fine-Tuned Models

Saving Fine-Tuned Models

Python โ€” Saving Different Fine-Tuning Methods
from peft import AutoPeftModelForCausalLM # Full fine-tuning: save complete model model.save_pretrained("./full_ft_model") # LoRA: save only adapter weights (~4-10 MB) peft_model.save_pretrained("./lora_weights") # Save tokenizer too tokenizer.save_pretrained("./full_ft_model") tokenizer.save_pretrained("./lora_weights")

Merging LoRA Adapters with Base Model

For deployment, you may want to merge LoRA weights into the base model. This creates a single model file without needing PEFT at inference.

Python โ€” Merging LoRA into Base Model
from peft import PeftModel, PeftConfig # Load the base model and LoRA weights base_model = AutoModelForCausalLM.from_pretrained("mistral-7b") peft_model = PeftModel.from_pretrained(base_model, "./lora_weights") # Merge: converts LoRA weights into base weights merged_model = peft_model.merge_and_unload() # Save merged model (can now be used with any inference framework) merged_model.save_pretrained("./merged_model") tokenizer.save_pretrained("./merged_model") # At inference, just load normally (no PEFT needed) from transformers import pipeline pipe = pipeline("text-generation", model="./merged_model")

Serving Fine-Tuned Models

Python โ€” REST API with FastAPI
from fastapi import FastAPI from transformers import pipeline import torch app = FastAPI() # Load model once at startup device = 0 if torch.cuda.is_available() else -1 pipe = pipeline( "text-generation", model="./fine_tuned_model", device=device, torch_dtype=torch.float16, ) @app.post("/predict") def predict(text: str, max_length: int = 100): outputs = pipe(text, max_length=max_length, num_return_sequences=1) return {"output": outputs[0]["generated_text"]} # Run: uvicorn app:app --host 0.0.0.0 --port 8000 # Curl: curl -X POST http://localhost:8000/predict -d '{"text": "Hello"}'

Quantization for Efficient Serving

Python โ€” ONNX Quantization
from optimum.onnxruntime import ORTModelForSequenceClassification # Convert to ONNX and quantize ort_model = ORTModelForSequenceClassification.from_pretrained( "./fine_tuned_model", from_transformers=True, file_name="model_quantized.onnx" ) # Save quantized model (much faster inference, 4x smaller) ort_model.save_pretrained("./quantized_model") # Use quantized model from transformers import pipeline pipe = pipeline("text-classification", model="./quantized_model")

Model Comparison and Versioning

Keep track of multiple fine-tuned versions:

Python โ€” Model Registry
import json from datetime import datetime model_registry = { "v1": { "date": "2024-01-15", "method": "full-ft", "dataset_size": 5000, "val_accuracy": 0.92, "path": "./models/ft_v1", }, "v2": { "date": "2024-01-20", "method": "lora", "dataset_size": 10000, "val_accuracy": 0.94, "path": "./models/ft_v2", } } with open("model_registry.json", "w") as f: json.dump(model_registry, f, indent=2) # Later, load the best version best_model_path = model_registry["v2"]["path"]

Troubleshooting Common Issues

Training Issues

RuntimeError: CUDA out of memory โ–ผ

Solutions (in order of effectiveness):

  1. Reduce batch size: per_device_train_batch_size=4 (from 16)
  2. Enable gradient checkpointing: gradient_checkpointing=True
  3. Reduce max sequence length: max_length=256 (from 512)
  4. Enable 8-bit optimization: Install bitsandbytes, use load_in_8bit=True
  5. Use LoRA: reduces memory by 50-70%
  6. Use QLoRA: reduces memory by 90%+
NaN loss or loss becomes 0 โ–ผ
  • NaN loss: Usually learning rate too high. Reduce by 10x.
  • Loss stays 0: Model has converged or collapsed. Try different initialization or check data.
  • Solutions: Use smaller learning rate, add warmup steps, check data for NaN values, try different random seed.
Model generates repetitive or nonsensical text โ–ผ
  • Mode collapse: Model learned a single "safe" output.
  • Solutions: Increase data diversity, lower learning rate, add temperature to decoding (temperature=0.7), check training examples for quality.

Data Issues

Data loading is very slow โ–ผ
  • Use Arrow format (.arrow) instead of JSON for faster loading
  • Set num_proc=4 in dataset.map() for parallel processing
  • Cache preprocessed data: dataset.save_to_disk("./cached_data")
  • Use smaller batches to see progress faster during development
Tokenizer adds special tokens I didn't expect โ–ผ
  • Check what special tokens your tokenizer has: tokenizer.special_tokens_map
  • Add custom tokens: tokenizer.add_special_tokens({'additional_special_tokens': ['', '']})
  • Resize model embeddings: model.resize_token_embeddings(len(tokenizer))

Inference Issues

Fine-tuned model doesn't work better than base model โ–ผ
  1. Check that you're loading the right model (fine-tuned version, not base)
  2. Evaluate on task-specific metrics, not generic ones
  3. Fine-tune longer: increase epochs to 5-10
  4. Add more data: may need 500+ examples for measurable improvement
  5. Check data quality: bad labels hurt more than no labels

Complete End-to-End Code Examples

Example 1: Fine-Tune BERT for Sentiment Classification

Python โ€” BERT Fine-Tuning Complete
#!/usr/bin/env python3 from datasets import load_dataset from transformers import AutoTokenizer, AutoModelForSequenceClassification from transformers import TrainingArguments, Trainer import torch # 1. Load dataset dataset = load_dataset("rotten_tomatoes") # 2. Tokenize tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased") def tokenize_function(examples): return tokenizer(examples["text"], padding="max_length", truncation=True, max_length=128) tokenized = dataset.map(tokenize_function, batched=True) # 3. Load model model = AutoModelForSequenceClassification.from_pretrained( "bert-base-uncased", num_labels=2 ) # 4. Training training_args = TrainingArguments( output_dir="./bert_sentiment", num_train_epochs=3, per_device_train_batch_size=32, per_device_eval_batch_size=64, learning_rate=2e-5, logging_steps=100, evaluation_strategy="epoch", save_strategy="epoch", load_best_model_at_end=True, metric_for_best_model="accuracy", ) trainer = Trainer( model=model, args=training_args, train_dataset=tokenized["train"], eval_dataset=tokenized["validation"], ) trainer.train() # 5. Inference from transformers import pipeline pipe = pipeline("text-classification", model="./bert_sentiment") print(pipe("This movie is amazing!"))

Example 2: LoRA Fine-Tune Mistral 7B

Python โ€” Mistral 7B with LoRA
#!/usr/bin/env python3 from peft import LoraConfig, get_peft_model from transformers import ( AutoModelForCausalLM, AutoTokenizer, Trainer, TrainingArguments, BitsAndBytesConfig, ) import torch # Load model with 4-bit quantization bnb_config = BitsAndBytesConfig( load_in_4bit=True, bnb_4bit_quant_type="nf4", bnb_4bit_compute_dtype=torch.bfloat16, ) model = AutoModelForCausalLM.from_pretrained( "mistralai/Mistral-7B-v0.1", quantization_config=bnb_config, device_map="auto", ) tokenizer = AutoTokenizer.from_pretrained("mistralai/Mistral-7B-v0.1") # Add LoRA adapters lora_config = LoraConfig( r=16, lora_alpha=32, target_modules=["q_proj", "v_proj"], lora_dropout=0.05, bias="none", task_type="CAUSAL_LM" ) model = get_peft_model(model, lora_config) print(model.print_trainable_parameters()) # Prepare data dataset = load_dataset("openwebtext", split="train[:1%]") def tokenize(examples): return tokenizer(examples["text"], truncation=True, max_length=512) tokenized = dataset.map(tokenize, batched=True, remove_columns=["text"]) # Train training_args = TrainingArguments( output_dir="./mistral_lora", num_train_epochs=1, per_device_train_batch_size=4, gradient_accumulation_steps=4, learning_rate=5e-5, warmup_steps=100, logging_steps=10, save_steps=500, fp16=True, ) trainer = Trainer( model=model, args=training_args, train_dataset=tokenized, ) trainer.train() model.save_pretrained("./mistral_lora")

Example 3: Instruction Tuning GPT-2

Python โ€” Instruction Tuning
#!/usr/bin/env python3 import json from datasets import Dataset from transformers import ( AutoModelForCausalLM, AutoTokenizer, Trainer, TrainingArguments, DataCollatorForLanguageModeling, ) # Create instruction dataset instructions = [ { "instruction": "Classify the sentiment", "input": "This product is great!", "output": "positive" }, { "instruction": "Classify the sentiment", "input": "Terrible experience", "output": "negative" }, ] def format_example(example): text = f"Instruction: {example['instruction']}\n" text += f"Input: {example['input']}\n" text += f"Output: {example['output']}<|endoftext|>" return {"text": text} formatted = [format_example(ex) for ex in instructions] dataset = Dataset.from_dict({"text": [ex["text"] for ex in formatted]}) # Tokenize tokenizer = AutoTokenizer.from_pretrained("gpt2") tokenizer.pad_token = tokenizer.eos_token def tokenize(examples): return tokenizer(examples["text"], truncation=True, max_length=512) tokenized = dataset.map(tokenize, batched=True) # Train model = AutoModelForCausalLM.from_pretrained("gpt2") training_args = TrainingArguments( output_dir="./instruction_gpt2", num_train_epochs=3, per_device_train_batch_size=8, learning_rate=5e-5, warmup_steps=100, save_steps=100, save_total_limit=2, ) trainer = Trainer( model=model, args=training_args, train_dataset=tokenized, data_collator=DataCollatorForLanguageModeling(tokenizer, mlm=False), ) trainer.train()

Hands-On Exercises

Exercise 1: Fine-Tune a Model on Your Own Data

Task: Collect or create 100-200 examples of text with labels (sentiment, intent, topic, etc.). Fine-tune a BERT or RoBERTa model on your data.

Learning goals: Understand the full fine-tuning pipeline, learn how to preprocess real data, evaluate on test set.

Deliverables: Working code, trained model, evaluation metrics on test set.

Exercise 2: Compare Full Fine-Tuning vs LoRA

Task: Fine-tune the same model using both full fine-tuning and LoRA on the same dataset. Compare final performance, training time, and model size.

Learning goals: Understand trade-offs between methods, see LoRA in action.

Deliverables: Comparison table with metrics and time/memory usage.

Exercise 3: Debug a Failing Fine-Tuning

Task: You're given a broken fine-tuning script. Fix the following issues: learning rate too high (NaN loss), data loading bug, incorrect model loading.

Learning goals: Practical debugging skills, common error recognition.

Exercise 4: Build a Custom Evaluation Metric

Task: For your domain (medical, legal, code, etc.), implement a custom evaluation metric beyond generic accuracy. Integrate it into HuggingFace Trainer.

Learning goals: Design metrics for specific use cases, integrate into training pipeline.

Interview Questions on Fine-Tuning

Why fine-tune instead of prompt engineering?

Fine-tuning adapts model internals to your task. Prompts are limited to context window. Fine-tuning gives better performance on specific tasks with custom data.

How do you prevent overfitting during fine-tuning?

Use validation set, early stopping, weight decay, dropout, data augmentation. Monitor train/val loss divergence. Fine-tune fewer epochs.

What's the difference between full fine-tuning and LoRA?

Full: updates all weights. LoRA: updates only low-rank adapters (0.1% of params). LoRA is more efficient but slightly lower quality (95% of full performance).

When would you use QLoRA over LoRA?

When memory is critical. QLoRA quantizes base model to 4-bit, reducing memory by 10x. Trade-off: slower training. Use if you only have consumer GPU.

How do you choose learning rate for fine-tuning?

Typically 1e-5 to 5e-4. Much lower than pre-training (5e-4 to 1e-3). Start with 2e-5, use learning rate finder or grid search if performance is poor.

What's the minimum data needed for fine-tuning?

50-100 examples can improve over base model. 500-2000 gives good results. 5000+ gives very good results. Depends on task complexity.

How do you know if a fine-tuned model is better than base?

Evaluate on task-specific metrics and domain-specific test set. Not on generic benchmarks. Compare inference speed and memory too.

Should you freeze earlier layers during fine-tuning?

Usually no. Unfrozen early layers can adapt to your domain better. Only freeze if you have very little data. Modern practice: fine-tune all layers.

Frequently Asked Questions

How long does fine-tuning take? ▼
Depends on method and hardware. Full fine-tuning on 7B model: 4-24 hours on H100/A100. LoRA: 1-4 hours. QLoRA: 2-8 hours on consumer GPU.
Can I fine-tune a model on CPU? ▼
Yes, but very slow (10-100x slower than GPU). Not practical for large models. Use Google Colab or cloud GPUs for practical fine-tuning.
How much data do I need? ▼
As little as 50 examples can help. But sweet spot is 500-2000 for good results. More data helps up to a point (diminishing returns around 10,000-50,000 examples).
Should I use LoRA or full fine-tuning? ▼
Start with LoRA. It's faster, cheaper, and good enough for most tasks (95% of full FT quality). Use full FT only if LoRA performance is insufficient and you have budget.
How do I know if I'm overfitting? ▼
Validation loss increases while training loss decreases. Model memorizes training examples instead of learning generalizable patterns. Use early stopping or regularization.
Can I fine-tune on multiple GPUs? ▼
Yes. HuggingFace Trainer supports DistributedDataParallel. Just set up your environment and Trainer handles the rest.
What's the difference between fine-tuning and transfer learning? ▼
Transfer learning is the broad concept. Fine-tuning is a specific transfer learning approach. All fine-tuning is transfer learning, but not all transfer learning is fine-tuning.
Can I save multiple adapters for different tasks? ▼
Yes with PEFT. Train different LoRA adapters, save them separately, load the right one at inference based on the task.
How do I serve a fine-tuned model in production? ▼
Merge LoRA weights into base model, then use standard inference (FastAPI, TorchServe, vLLM, etc.). Or quantize for faster inference.
Is fine-tuning still relevant with in-context learning? ▼
Yes. Large context windows help, but fine-tuning still beats few-shot for specific domains. Hybrid: few-shot + retrieval + fine-tuning often wins.

Summary: Key Takeaways

Fine-tuning is transfer learning

Adapt pre-trained models to your task. Requires 10-100x less data than training from scratch.

LoRA is the practical default

Fine-tune with 0.1-1% of parameters. 95% of full FT performance at 1/10 the cost. PEFT library makes it easy.

Data quality matters most

100 perfect examples beat 10,000 noisy ones. Spend time cleaning, deduplicating, validating.

Use validation sets always

Monitor validation metrics during training. Detect overfitting early with early stopping.

Start small and iterate

Begin with 100-200 examples and LoRA. Increase data and switch methods if needed.

Lower learning rates for fine-tuning

Use 1e-5 to 5e-4, not 5e-4 to 1e-3. Pre-trained weights are delicate.

The Fine-Tuning Workflow

1. Prepare Data
Clean, split, validate (100-2000 examples)
โ†“
2. Choose Method
LoRA (default) or full FT or QLoRA
โ†“
3. Setup Training
Learning rate, batch size, epochs
โ†“
4. Monitor & Evaluate
Watch for overfitting, use early stopping
โ†“
5. Deploy
Save, merge (if LoRA), serve with FastAPI/vLLM

Further Resources

Libraries and Tools

  • HuggingFace Transformers: https://github.com/huggingface/transformers โ€” industry standard library
  • PEFT: https://github.com/huggingface/peft โ€” parameter efficient fine-tuning methods
  • TRL: https://github.com/huggingface/trl โ€” RLHF, DPO, and alignment
  • vLLM: https://github.com/lm-sys/vllm โ€” fast LLM inference
  • Optuna: https://optuna.org โ€” hyperparameter optimization
  • Weights & Biases: https://wandb.ai โ€” experiment tracking

Papers to Read

Practical Guides

  • HuggingFace Course: https://huggingface.co/course โ€” free comprehensive course
  • LoRA Fine-Tuning Guide: https://huggingface.co/docs/peft/conceptual_guides/lora
  • Llama 2 Fine-Tuning: Meta's official guide and examples
  • HuggingFace Blog: Regular posts on fine-tuning techniques and best practices

Communities

  • HuggingFace Forums: Active community, questions answered by maintainers
  • r/MachineLearning: Reddit community with fine-tuning discussions
  • Hugging Face Slack: Real-time chat with researchers and practitioners
  • Open LLM Leaderboard: See what fine-tuning techniques lead to SOTA

Next Steps

Now that you understand fine-tuning fundamentals:

  • Pick a task and dataset that interests you
  • Start with LoRA fine-tuning on a small model (7B-13B)
  • Experiment with hyperparameters using the techniques here
  • Share results and get feedback from the community
  • When comfortable, explore RLHF/DPO for alignment