Most teams asking for a fine-tune in September 2026 are asking the wrong vendor, and often the wrong technique. OpenAI told developers on May 7, 2026 that self-serve fine-tuning is closing: organizations that have never run a job cannot start one, organizations that have not called a fine-tuned model in 60 days lost job creation on July 2, and active customers lose the ability to create new jobs on January 6, 2027. Inference on a model you already trained survives only until that base snapshot is deprecated. Several fine-tuned snapshots have an earlier, explicit shutdown on October 23, 2026. This AI fine-tuning guide 2026 when you need it is that decision test. The April technique split is linked below; it predates the wind-down.
Read Fine-Tuning vs RAG vs Prompt Engineering if you still need the technique definitions. What follows assumes you know the difference between a prompt, a retrieval index, and a gradient step. The question that matters this quarter is narrower: given a failure you can measure, is a training job still the move, and if it is, which console will still accept the file?

The dates that retired the old advice
OpenAI’s deprecations page, as fetched September 28, 2026, states the wind-down in three rows. On May 7, 2026, creating a fine-tuning job became unavailable to organizations that had not previously run fine-tuning. On July 2, 2026, job creation closed for organizations that had not run inference on a fine-tuned model in the prior 60 days. On January 6, 2027, active existing customers can no longer create new jobs. The same section says inference continues until the underlying base model is deprecated. That sentence is not a blank check. A later block on the same page schedules explicit shutdowns.
The clean fine-tuned-snapshot table lists October 23, 2026 as the shutdown date for
ft-gpt-3.5-turbo
,
ft-gpt-4
,
ft-gpt-4.1-nano-2025-04-14
,
ft-babbage-002
, and
ft-davinci-002
. The same dated block also lists
ft-o4-mini-2025-04-16
and
o4-mini-2025-04-16
, with
gpt-5.6-terra
as the substitute, and lists
gpt-4.1-nano
and
gpt-4.1-nano-2025-04-14
with
gpt-5.6-luna
as the substitute. If your production model ID is in that block, “until the base is deprecated” has already been given a date. If it is not, you still cannot start a replacement job after January 6, 2027.
The model optimization guide carries the same wind-down banner and, underneath it, the method table that is still published. Supervised fine-tuning and direct preference optimization are listed for
gpt-4.1-2025-04-14
,
gpt-4.1-mini-2025-04-14
, and
gpt-4.1-nano-2025-04-14
. Vision fine-tuning is listed for
gpt-4o-2024-08-06
only. Reinforcement fine-tuning is listed for
o4-mini-2025-04-16
only. GPT-5.6 appears on the deprecations page as a recommended replacement, not as a row in that method table. As fetched, there is no self-serve fine-tune path onto the model OpenAI is telling you to migrate to.
What a training job changes, and what it cannot
A fine-tune updates weights so a narrow behavior becomes more likely. OpenAI’s own list of reasons to do it is specific: more demonstrations than fit in one context window, shorter prompts once the examples move into the weights, training on data you do not want to resend on every request, and teaching a smaller model a task where the large model is not cost-effective. None of those reasons is “the model does not know our docs.” Retrieval is the tool for facts that change or that were never in the pretraining mix. The production write-up is RAG isn’t dead. If the failure is a missing clause in a contract the model has never seen, a gradient step on last quarter’s tickets will not insert it.
It also will not invent an evaluation. The optimization guide still puts evals first, then prompting, then fine-tuning as an optional third step for some use cases. A training job without a held-out set measures memorization. OpenAI’s best-practices page is blunt about the ceiling: if the people who wrote the labels only agreed on 70 percent of extracted spans, the model is unlikely to beat that agreement. The dataset is the spec. A sloppy spec produces a sloppy checkpoint, faster than a sloppy prompt does, and harder to roll back.
Lab-side preference training is a different object. RLHF, as the labs run it, is how a base model becomes an assistant. Your supervised fine-tune on 400 support macros is not that pipeline, and calling it RLHF in a roadmap deck will get the data format wrong. Direct preference optimization on OpenAI’s platform is a customer job with preferred and rejected completions. It is not a license to rerun the lab’s reward-model stack.
Five questions before anyone opens a training console
Run these in order. The first “no” that is actually a different technique should stop the project. The April comparison still describes the techniques. This sequence adds the eval gate and the platform gate that comparison did not have, because in April the platform gate was not the binding constraint.
- Is the failure missing or stale facts? If yes, stop. Build or fix retrieval. A fine-tune memorizes the facts in the file you upload and then serves them as if they were still true.
- Is the failure format, tone, or a schema a prompt can lock? If a JSON schema, a few-shot block, or a stricter instruction moves the eval, you do not have a weights problem. You have a prompt that has not been finished. Prompting still pays precisely on this class of failure.
- Does a held-out eval show that prompt plus retrieval has plateaued? If you cannot answer, you are not ready to train. Write the eval. Score the current stack. A training job with no baseline is a budget line, not an experiment.
- Do you have the right kind of signal for the method? Correct completions for supervised fine-tuning. Preferred and rejected pairs for DPO. A programmatic grader that qualified people would agree with, for reinforcement fine-tuning. “We will label it during the job” is not a dataset.
- Can you create the job on a platform that will still be there when you need to retrain? This is the 2026 question. A checkpoint you cannot refresh is a depreciating asset with a published sunset.
Four of those five exits do not produce a training file. That is the point. The teams that should fine-tune are the ones who survive all five, not the ones who liked a conference talk.
Two requests, only one of which should become a job
Take a support classifier that must emit one of 40 product codes and nothing else. The base model writes a paragraph, then a code, then an apology. A JSON schema cuts the apology and still misses 6 codes the taxonomy named last year. Few-shot examples of those 6 codes blow the prompt past the latency budget, and a held-out set of 200 tickets, labeled by two people who agreed on 94 percent of rows, still fails. That request survives the five questions, provided the organization can still open a console. The file is SFT-shaped: ticket in, code out. The eval exists. The labels agree. The only remaining argument is which base, and whether you will be allowed to retrain it in March.
Now take the request that arrives in the same standup: “fine-tune the model on our help center so it stops inventing refund windows.” The failure is a fact. The refund window changed in June. A training file built from the old articles will teach the old window more confidently. The fix is the index behind the production RAG piece, plus a prompt that must quote the retrieved clause. No example count repairs that. If someone on the thread says the fine-tune will “learn the docs,” they have confused a snapshot with a subscription.
A third shape shows up often enough to name. The team has 30 beautiful examples and no disagreement data. That is a prompt. Put the 30 in the system message, or in a retrieval set of exemplars, and measure. OpenAI’s current best-practices page does not bless a minimum of 30, and it does warn that a small high-quality set beats a large messy one only after you have decided the set is the spec. Thirty rows that all come from one senior agent are that agent’s habits, not the queue. If the night shift handles a different mix, the checkpoint will not know. Collect the disagreement before you collect the epoch.
When the prompt is the spend
OpenAI’s optimization guide says the prompt-engineering loop may be all a use case needs, and that you should include context the model does not have, state the output you want, and show a few correct examples. That is few-shot learning inside the request, not a fine-tune. It is also reversible in an afternoon. A training job is not.
Stay on the prompt when the task is still moving. Classification taxonomies that change every sprint, tone guides the brand team rewrites monthly, and tool schemas that gained three fields last week are all bad fine-tune targets. You would be baking a version into weights and then paying to bake the next version. Google’s supervised-tuning docs make the same cut from the other direction: tune when the behavior is hard to articulate in a prompt, or when you want to drop the few-shot examples to shorten the context. If you can still say the rule in a system message and the eval passes, you have not earned the job.
Cost belongs in this decision, but not as a vibe. A fine-tune that lets you delete 2,000 tokens of examples from every request can pay for itself at volume. A fine-tune that does not shorten the prompt, and that you will retrain twice before January, does not. The site’s 2026 pricing guide is the place to price the base call. Do not invent a training-token rate from a blog that is not the vendor page. OpenAI points training and usage billing at its pricing page; this guide does not copy a number that was not on the pages fetched today.
When retrieval is the spend
Use retrieval when the answer has to track a corpus you do not control: policy PDFs, ticket history, a catalog, a codebase that moved yesterday. Fine-tuning does not subscribe to that corpus. It snapshots whatever you put in the JSONL. The next edit to the source document is invisible to the checkpoint until you rebuild the file and, if the console still lets you, run another job.
The failure mode to watch is the hybrid that is not a hybrid. Teams upload the docs as fine-tune examples, call it “our knowledge,” and then discover the model cites a paragraph that was deleted in March. If the eval includes a freshness slice and that slice fails, the fix is the index, not another epoch. Retrieval plus a prompt is also the path that does not care whether OpenAI’s fine-tune console is open. That is not a consolation prize in September 2026. It is the path with no January cutoff.
When a fine-tune still pays
The cases that survive the five questions are narrow and repetitive. Google’s tuning overview names them: classification where the label is a specific phrase and the base model will not stop explaining itself; summarization with a format you cannot reliably describe, such as replacing speaker names with stable placeholders; extractive answers that must be a substring of a supplied context; a chat persona that prompting will not hold. OpenAI’s method table points at the same neighborhood: classification, nuanced translation, a specific output format, instruction-following failures the prompt does not fix, and, for DPO, summaries and tone where you can rank two answers but cannot write the perfect one.
Volume matters only after the eval moves. If you send the same extraction task hundreds of thousands of times a month, and a tuned smaller model matches the prompted large model on your held-out set, the token delta is real money. When a 7B model beats a 70B is the same shape of argument on open weights: the win is task-fit, not parameter count. A one-off report, a workflow that changes weekly, or a task where the base model already scores at the top of your rubric is not that case. OpenAI’s reinforcement-fine-tuning guide says the quiet part: if the model you want to tune already scores at the floor or the ceiling of the grader, RFT has nothing to reinforce.
Pick the method from the file you actually have
Do not pick a method because the acronym is newer. Pick it because the file in front of you matches the schema. As fetched on September 28, 2026, the published OpenAI table is:
| Method | What the file contains | Listed base, as fetched | Stop if |
|---|---|---|---|
| Supervised fine-tuning | Prompt plus a correct response |
gpt-4.1-2025-04-14
, mini, nano | You cannot write the correct answer, or the labels disagree |
| Direct preference optimization | Prompt, preferred output, non-preferred output. Text only. | Same three GPT-4.1 snapshots | You need a single canonical answer, not a ranking |
| Vision fine-tuning | Image input plus the desired response |
gpt-4o-2024-08-06
only | The task is image generation. This is understanding, not drawing |
| Reinforcement fine-tuning | A prompt plus a grader, not a fixed answer |
o4-mini-2025-04-16
only | Experts do not converge, or the base already sits at the grader’s floor or ceiling |
| Gemini supervised tuning | Labeled examples, hundreds of them in Google’s own framing | Gemini 2.5 Pro, 2.5 Flash, 2.5 Flash-Lite | You need a Covered Service SLA. Google says this tuning is not one |
The SFT guide, the DPO guide, and the RFT guide each repeat the wind-down banner above the method. Read the banner before you read the code sample. A copy-pasteable job create call is not evidence that your organization can submit it.
DPO is the one people mis-file under RLHF. On this platform it is a preference-pair job on a GPT-4.1 snapshot. OpenAI notes that running SFT on the preferred responses first, then a DPO job, can improve alignment. That is a two-job sequence, not a reason to skip the eval between them. RFT is the one people start without a grader. The guide’s precondition is stricter than a rubric in a doc: qualified people, given only what the model sees, have to converge on the same answers, and the grader you ship has to be able to score that agreement. Theorem-style tasks, tests that pass or fail, schema validation. Not “sounds more helpful.”
Where a job can still be created
Three doors are still documented. They are not interchangeable, and one of them is already closing.
OpenAI, if you are already inside. New organizations have been out since May 7. Organizations that went quiet on fine-tuned inference have been out since July 2. Everyone else has until January 6, 2027 to create jobs, and a subset of snapshots die on October 23, 2026 regardless. If you are inside, the useful work this month is an inventory: model ID, last inference date, whether that ID is in the October 23 block, and whether you can rebuild the dataset without the console. A checkpoint you cannot retrain is a migration project, not an asset.
Gemini supervised fine-tuning, on the models the page actually lists. Google’s supervised fine-tuning document, last updated May 5, 2026 and refetched today, lists Gemini 2.5 Pro, Gemini 2.5 Flash, and Gemini 2.5 Flash-Lite. The limitations tables also cover Gemini 2.0 Flash and Flash-Lite. Adapter sizes on 2.5 Flash and Flash-Lite are 1, 2, 4, 8, and 16; on 2.5 Pro they are 1, 2, 4, and 8. Training examples are capped at 131,072 input and output tokens. The training file cap is 1 GB of JSONL, up to 10 million text examples or 300,000 multimodal examples. Supervised fine-tuning is not a Covered Service and is excluded from the SLO of any SLA. The same page says inference pricing follows the stable base version. It does not, on the page fetched today, list a Gemini 3 model as a tuning base. Do not promote a 2.5 job into a Gemini 3 claim.
Two Gemini-specific traps are in that document and worth reading before you celebrate the open door. Controlled generation at inference time on a tuned Gemini model can drop quality, because tuning did not apply it; the recommended substitute is to bake the structure into the training targets. And for thinking models, Google suggests setting the thinking budget to off or to its lowest value for the tuned task, because the tuned model is trained to skip the thinking trace. A tuned 2.5 checkpoint is not a drop-in for a thinking-on production prompt.
Open weights, if you can run the adapter yourself. As fetched today, Anthropic’s models overview does not list a fine-tuning feature, and the conventional Claude fine-tuning doc paths return 404. That is an absence in the published docs, not a sworn statement that no private program exists. It is enough to take Claude API off the list of self-serve training consoles. If the behavior you need has to live in weights and neither OpenAI nor a Gemini 2.5 tune fits, the remaining documented path is an open-weight model plus a parameter-efficient update. LoRA, the 2021 method paper, freezes the base and trains low-rank update matrices instead of every weight (Hu et al., arXiv:2106.09685). Whether that is cheaper than a prompted frontier model depends on your hardware and your eval, which is the argument in when local models beat cloud and in open source versus closed. The network’s June note on the PEFT revival, Fine-Tuning Is Back, is the opinionated version of that door. Treat it as an argument, not as a vendor spec.
The data bar, without a magic example count
Older cookbook pages told people to start around 50 to 100 demonstrations on GPT-3.5-class models, with a hard floor of 10. That number is not what the current best-practices page says, and this guide will not launder a 2023 heuristic into a 2026 requirement. What the current page does say is more useful. Quality beats quantity when you have to choose. If the model still fails a slice, add examples that show that slice done correctly. If 60 percent of training replies are refusals and only 5 percent of live traffic should refuse, you will get refusals. If a compliment in the target mentions a trait that was not in the prompt, the model learns to invent traits. Every example has to be in the format you will send at inference, or you are training a different API than the one you will call.
The quantity test they do publish is a doubling probe: fine-tune on the full set, fine-tune on half, and read the gap. They expect a similar gain each time you double, until the labels, not the count, are the limit. Google’s tuning introduction is the other official quantity line worth keeping: supervised tuning is framed around hundreds of labeled examples that teach the model to mimic a task. “Hundreds” is not 12 rows exported from a spreadsheet on Friday.
Hold out a test file the training job never sees. OpenAI will report stats on a validation file you attach to the job. That is necessary and not sufficient. A validation split cut from the same labeling session will not catch a taxonomy change, a new document type, or the freshness slice that retrieval was supposed to own. If you cannot name the slice that would make you throw the checkpoint away, you do not have an eval. You have a hope.
What to do this week
If a fine-tuned OpenAI model is in production, export the model ID and the last inference timestamp before you do anything else. Check it against the October 23, 2026 snapshot list and against the July 2 activity rule. A model you have not called in 60 days may already be unable to spawn a successor job, even though inference on the old checkpoint still works. Put a migration owner on any ID in the October 23 block. The substitute column on that page points at GPT-5.6 tiers, which are not listed as fine-tune bases. Plan on a prompt, a retrieval index, a Gemini 2.5 tune, or an open-weight adapter. Do not plan on “we will just re-tune GPT-5.6 in the same console.”
If you are about to start a project, write the eval this week and score prompt plus retrieval before anyone requests a training budget. Most of those scores will end the conversation. The ones that do not should arrive with a file that matches one row of the method table, a platform that can still accept the job, and a date when that platform stops accepting it. That date is January 6, 2027 for new OpenAI jobs, earlier for the snapshots already scheduled, and unpublished for a Gemini 3 tune because that tune is not in the document. The decision is a calendar problem now. Treat it like one.
FAQ
Can I fine-tune GPT-5.6?
Not on the method table fetched September 28, 2026. Supervised fine-tuning and DPO are listed for three GPT-4.1 snapshots. Vision fine-tuning is listed for
gpt-4o-2024-08-06
. Reinforcement fine-tuning is listed for
o4-mini-2025-04-16
. GPT-5.6 is the recommended substitute for models that are shutting down. It is not listed as a fine-tune base. Recheck the optimization guide before you tell a stakeholder otherwise. The table can move. As of this fetch, it has not moved onto GPT-5.6.
Does fine-tuning replace RAG?
No. Fine-tuning changes how the model behaves on a pattern. Retrieval supplies facts that were not in the weights, including facts that changed after the training file was built. If the eval failure is a missing or stale fact, another epoch will not fix it. Use the retrieval architecture, and keep fine-tuning for format, classification, and behaviors you cannot hold with a prompt.
What happens to existing OpenAI fine-tunes after January 6, 2027?
January 6, 2027 is the date OpenAI’s deprecations page gives for the end of new job creation by active existing customers. It is not, by itself, the date inference stops. The same page says inference on fine-tuned models continues until the underlying base model is deprecated. Separately, specific fine-tuned snapshots, including
ft-gpt-4.1-nano-2025-04-14
and
ft-o4-mini-2025-04-16
, are listed for shutdown on October 23, 2026. Read your model ID against both sentences. One is a job-creation cutoff. The other is a per-snapshot shutdown.
Is DPO the same thing as RLHF?
No. RLHF, in the lab pipeline, trains a reward model on preferences and then updates the policy against that reward. DPO on OpenAI’s platform is a customer fine-tune that takes a preferred completion and a non-preferred completion for the same prompt, currently listed only for text, on GPT-4.1 snapshots. Related technique, different job, different data file. If your spreadsheet has one “good answer” column, you have an SFT file, not a DPO file.
Should a new team start a fine-tune in September 2026?
Only if all five questions above come back yes, and only on a console your organization can actually open. A new OpenAI organization has not been able to create fine-tuning jobs since May 7, 2026. A Gemini 2.5 supervised tune is still documented, with the SLA exclusion and the model list above. An open-weight LoRA is still a method, not a console. If any of the first four questions is no, do not start. Spend the week on the eval and the prompt. The training job will still be the wrong tool in October.
Related reading
- Fine-Tuning vs RAG vs Prompt Engineering — the April technique split this guide does not repeat.
- Prompt engineering isn’t dead — when the failure is still a prompt.
- RAG in production — when the failure is facts.
- RLHF, plain English — lab preference training versus your customer job.
- Which AI tool should I use in 2026 — the adjacent decision, one level up from the training console.
- Local models, when hardware wins — the open-weight door if both consoles are wrong.