โ† Back to list

SFT โ€” Supervised Fine-Tuning: AI Ko โ€œSikhayaโ€ Kaise Jaata Hai? ๐ŸŽ“

GPT ne sab kuch padha. Billions of web pages. Books. Code. Conversations.

Dhanashree ยท 2026-06-07 06:09 ยท 10 claps ยท 2.8 min read
#sft #ai #artificial-intelligence #technology #tech
Open on Medium โ†—
Wiki topics: FT ยท Fine-tuning & Adaptation AI ยท AI ยท General

SFT โ€” Supervised Fine-Tuning: AI Ko โ€œSikhayaโ€ Kaise Jaata Hai? ๐ŸŽ“

GPT ne sab kuch padha. Billions of web pages. Books. Code. Conversations.

Phir bhi โ€” ek raw, pre-trained model se agar seedha poochho kuch bhi:

It rambles. It goes off-track. It ignores what you actually asked.

โ€œBahut padha-likha hai, par saleeqa nahi.โ€ ๐Ÿ˜…

Thatโ€™s exactly where SFT โ€” Supervised Fine-Tuning comes in.

๐Ÿง  Pehle Samjho โ€” Pre-training vs Fine-Tuning

Think of it like this.

Pre-training = A student who reads every book in the library. All genres. All subjects. No filter.

Fine-Tuning = That same student now joins a specific coaching class โ€” with a teacher, structured lessons, and targeted feedback.

Same brain. Different direction.

โ€œLibrary mein sab kuch tha. Coaching ne focus diya.โ€ ๐ŸŽฏ

SFT is that coaching class for AI models.

๐Ÿ“‹ So What Exactly Is SFT?

Supervised Fine-Tuning is a training technique where you take a pre-trained model and train it further on a curated, labeled dataset โ€” with correct input-output pairs.

The format is simple:

  • Input: A question or instruction
  • Output: The ideal, human-written response

Example:

Input: โ€œExplain recursion in simple terms.โ€ Output: โ€œRecursion is when a function calls itself to solve a smaller version of the same problemโ€ฆโ€

The model sees thousands of such pairs. It learns: โ€œJab aisa poochha jaye, toh aisa jawab dena chahiye.โ€

๐Ÿ‘‰ Itโ€™s not just about knowledge anymore. Itโ€™s about behaviour.

๐Ÿ‘ฉโ€๐Ÿซ Who Writes These Correct Answers?

Humans. Real people. Called human annotators.

They are given guidelines โ€” be helpful, be honest, be safe โ€” and they write ideal responses to hundreds of different prompts.

This labeled data becomes the training signal for SFT.

โ€œAI ko sikhane ke liye pehle insaano ko likhna padta hai. Ironic, na?โ€ ๐Ÿ˜‚

The quality of this data is everything. Garbage in, garbage out โ€” as they say.

โœ”๏ธ Good annotators = Model that follows instructions well

โŒ Inconsistent annotations = Model that behaves unpredictably

โšก Why SFT Matters โ€” The Before vs After

Without SFT, a model might:

  1. ๐Ÿ”ด Complete your sentence instead of answering your question
  2. ๐Ÿ”ด Ignore instructions and go off on a tangent
  3. ๐Ÿ”ด Give a technically correct but totally unhelpful response

With SFT, the model learns to:

  1. โœ… Follow instructions precisely
  2. โœ… Respond in the right format and tone
  3. โœ… Stay on topic โ€” and actually be useful

โ€œPre-training ne dimag diya. SFT ne tameez sikhayi.โ€ ๐Ÿ™

๐Ÿ”— SFT in the Bigger Picture โ€” RLHF ka Bada Bhai

If youโ€™ve heard of RLHF (Reinforcement Learning from Human Feedback) โ€” SFT is Step 1 of that pipeline.

The full journey looks like:

Pre-training โ†’ SFT โ†’ Reward Modelling โ†’ RLHF

SFT gives the model a strong, instruction-following foundation.

Then RLHF refines it further โ€” using human preferences to reward better responses and penalise bad ones.

โ€œSFT ne base banaya. RLHF ne usse aur sharpen kiya.โ€ ๐Ÿ”ช

ChatGPT, Claude, Gemini โ€” sab isi pipeline se guzre hain.

โš ๏ธ The Catch โ€” SFT is Not Perfect

SFT is powerful. But it has limits.

It overfits to the style of annotators. If annotators always write formal responses, the model gets stiff.

It can be expensive. High-quality labeled data doesnโ€™t come cheap โ€” or fast.

It doesnโ€™t teach the model to reason. It teaches it to imitate correct answers. Big difference.

โ€œSirf copy karna seekha โ€” samajhna nahi. Exam mein yeh kaam nahi aata.โ€ ๐Ÿ˜ฌ

Thatโ€™s why SFT is always a starting point โ€” not the finish line.

๐Ÿ”ฎ Whatโ€™s Next After SFT?

The field is moving fast.

Researchers are now exploring:

  • RLAIF โ€” AI giving feedback instead of humans (cheaper, faster)
  • DPO (Direct Preference Optimisation) โ€” skipping reward models entirely
  • Synthetic data for SFT โ€” using AI-generated labeled data to reduce human effort

โ€œEk din AI hi AI ko sikhayega. Aur hum log sirf dekhenge.โ€ ๐Ÿคฏ

๐ŸŽฏ Final Thoughts

SFT is the bridge between a model that knows things and a model that actually helps you.

Itโ€™s not glamorous. Itโ€™s not the flashy part of AI. But without it, every LLM would be a brilliant mess โ€” like a topper who canโ€™t answer a straight question.

โ€œKnowledge toh tha. SFT ne use kaam ka banaya.โ€ ๐Ÿ’ก

Never heard of SFT before today? Drop a ๐Ÿ™‹ in the comments. And if you want more AI concepts broken down โ€” desi style โ€” hit follow! Bahut kuch samjhana baaki hai. ๐Ÿ‘‡


๋ฉ”ํƒ€๋ฐ์ดํ„ฐ
post_id
fdd6052f7dc4
slug
sft-supervised-fine-tuning-ai-ko-sikhaya-kaise-jaata-hai-fdd6052f7dc4
url
https://medium.com/@dhanashreeA/sft-supervised-fine-tuning-ai-ko-sikhaya-kaise-jaata-hai-fdd6052f7dc4
canonical_url
https://medium.com/@dhanashreeA/sft-supervised-fine-tuning-ai-ko-sikhaya-kaise-jaata-hai-fdd6052f7dc4
author_url
https://medium.com/@dhanashreeA
status
ok
fetched_at
2026-06-13 09:11:36