SFT โ Supervised Fine-Tuning: AI Ko โSikhayaโ Kaise Jaata Hai? ๐
GPT ne sab kuch padha. Billions of web pages. Books. Code. Conversations.
SFT โ Supervised Fine-Tuning: AI Ko โSikhayaโ Kaise Jaata Hai? ๐
GPT ne sab kuch padha. Billions of web pages. Books. Code. Conversations.
Phir bhi โ ek raw, pre-trained model se agar seedha poochho kuch bhi:
It rambles. It goes off-track. It ignores what you actually asked.
โBahut padha-likha hai, par saleeqa nahi.โ ๐
Thatโs exactly where SFT โ Supervised Fine-Tuning comes in.
๐ง Pehle Samjho โ Pre-training vs Fine-Tuning
Think of it like this.
Pre-training = A student who reads every book in the library. All genres. All subjects. No filter.
Fine-Tuning = That same student now joins a specific coaching class โ with a teacher, structured lessons, and targeted feedback.
Same brain. Different direction.
โLibrary mein sab kuch tha. Coaching ne focus diya.โ ๐ฏ
SFT is that coaching class for AI models.
๐ So What Exactly Is SFT?
Supervised Fine-Tuning is a training technique where you take a pre-trained model and train it further on a curated, labeled dataset โ with correct input-output pairs.
The format is simple:
- Input: A question or instruction
- Output: The ideal, human-written response
Example:
Input: โExplain recursion in simple terms.โ Output: โRecursion is when a function calls itself to solve a smaller version of the same problemโฆโ
The model sees thousands of such pairs. It learns: โJab aisa poochha jaye, toh aisa jawab dena chahiye.โ
๐ Itโs not just about knowledge anymore. Itโs about behaviour.
๐ฉโ๐ซ Who Writes These Correct Answers?
Humans. Real people. Called human annotators.
They are given guidelines โ be helpful, be honest, be safe โ and they write ideal responses to hundreds of different prompts.
This labeled data becomes the training signal for SFT.
โAI ko sikhane ke liye pehle insaano ko likhna padta hai. Ironic, na?โ ๐
The quality of this data is everything. Garbage in, garbage out โ as they say.
โ๏ธ Good annotators = Model that follows instructions well
โ Inconsistent annotations = Model that behaves unpredictably
โก Why SFT Matters โ The Before vs After
Without SFT, a model might:
- ๐ด Complete your sentence instead of answering your question
- ๐ด Ignore instructions and go off on a tangent
- ๐ด Give a technically correct but totally unhelpful response
With SFT, the model learns to:
- โ Follow instructions precisely
- โ Respond in the right format and tone
- โ Stay on topic โ and actually be useful
โPre-training ne dimag diya. SFT ne tameez sikhayi.โ ๐
๐ SFT in the Bigger Picture โ RLHF ka Bada Bhai
If youโve heard of RLHF (Reinforcement Learning from Human Feedback) โ SFT is Step 1 of that pipeline.
The full journey looks like:
Pre-training โ SFT โ Reward Modelling โ RLHF
SFT gives the model a strong, instruction-following foundation.
Then RLHF refines it further โ using human preferences to reward better responses and penalise bad ones.
โSFT ne base banaya. RLHF ne usse aur sharpen kiya.โ ๐ช
ChatGPT, Claude, Gemini โ sab isi pipeline se guzre hain.
โ ๏ธ The Catch โ SFT is Not Perfect
SFT is powerful. But it has limits.
It overfits to the style of annotators. If annotators always write formal responses, the model gets stiff.
It can be expensive. High-quality labeled data doesnโt come cheap โ or fast.
It doesnโt teach the model to reason. It teaches it to imitate correct answers. Big difference.
โSirf copy karna seekha โ samajhna nahi. Exam mein yeh kaam nahi aata.โ ๐ฌ
Thatโs why SFT is always a starting point โ not the finish line.
๐ฎ Whatโs Next After SFT?
The field is moving fast.
Researchers are now exploring:
- RLAIF โ AI giving feedback instead of humans (cheaper, faster)
- DPO (Direct Preference Optimisation) โ skipping reward models entirely
- Synthetic data for SFT โ using AI-generated labeled data to reduce human effort
โEk din AI hi AI ko sikhayega. Aur hum log sirf dekhenge.โ ๐คฏ
๐ฏ Final Thoughts
SFT is the bridge between a model that knows things and a model that actually helps you.
Itโs not glamorous. Itโs not the flashy part of AI. But without it, every LLM would be a brilliant mess โ like a topper who canโt answer a straight question.
โKnowledge toh tha. SFT ne use kaam ka banaya.โ ๐ก
Never heard of SFT before today? Drop a ๐ in the comments. And if you want more AI concepts broken down โ desi style โ hit follow! Bahut kuch samjhana baaki hai. ๐
๋ฉํ๋ฐ์ดํฐ
- post_id
- fdd6052f7dc4
- slug
- sft-supervised-fine-tuning-ai-ko-sikhaya-kaise-jaata-hai-fdd6052f7dc4
- url
- https://medium.com/@dhanashreeA/sft-supervised-fine-tuning-ai-ko-sikhaya-kaise-jaata-hai-fdd6052f7dc4
- canonical_url
- https://medium.com/@dhanashreeA/sft-supervised-fine-tuning-ai-ko-sikhaya-kaise-jaata-hai-fdd6052f7dc4
- author_url
- https://medium.com/@dhanashreeA
- status
- ok
- fetched_at
- 2026-06-13 09:11:36