Uncategorized

RAG or Fine-Tuning? Choosing an Approach for a Knowledge-Based AI Product

Quick answer

Start with prompting and a test set. Choose retrieval-augmented generation (RAG) when answers depend on changing documents or traceable evidence. Consider fine-tuning when the model receives the right information but repeatedly fails to follow the required behavior or format. Combine them only when testing shows both problems. These approaches address different needs; they are not mandatory stages of product maturity. (developers.openai.com)

1. Separate missing knowledge from inconsistent behavior

For a knowledge-based AI product, the first question is not “Which technology is more advanced?” It is “Why does the current system fail?”

Three approaches change different parts of the system:

  • Prompting supplies instructions, examples, and reference material in the model’s input.
  • RAG finds relevant material from an external collection and supplies it when answering a question.
  • Supervised fine-tuning trains a model on example inputs and desired outputs to shape its behavior. (developers.openai.com)

The original RAG research combined a generative model with an external document index. It explicitly addressed limitations around accessing knowledge, updating it, and establishing where information came from. Its benchmark results do not establish how well your company’s policy assistant will perform. (arxiv.org)

Use a simple diagnostic experiment: give the model the correct, complete evidence directly.

If the answer improves, investigate retrieval or missing context. If it still ignores exceptions, invents unsupported conclusions, or follows the wrong response pattern, investigate instructions and behavior. OpenAI’s optimization guidance makes the same distinction between supplying the wrong context and mishandling the right context. (developers.openai.com)

Practical check: save the question, supplied evidence, output, and reviewer’s explanation for each failure. “Wrong answer” alone is not a useful engineering diagnosis.

2. Use a decision matrix before choosing an architecture

The following matrix is a starting heuristic, not a performance guarantee. It translates the documented distinction between context optimization and behavior optimization into product decisions. (developers.openai.com)

Product requirement Start with What to check
A short, stable reference document Prompting with reference text Can the model answer all required cases from that input?
Frequently revised policies or manuals RAG plus prompting Are the applicable versions searchable and retrieved?
Answers with inspectable supporting evidence RAG or directly supplied sources Does each cited passage actually support the claim?
Consistent classification or specialized response behavior Prompting; then consider fine-tuning Does the failure persist with correct context?
Changing knowledge and persistent behavior failures RAG plus possible fine-tuning Can each component’s contribution be measured separately?
A deterministic eligibility rule or calculation Ordinary software, possibly with an AI explanation layer Can the decision be computed without generation?

For an early product, write acceptance criteria before selecting infrastructure. A policy assistant might need to identify the applicable policy, preserve exceptions, cite evidence, and decline unsupported questions.

Avoid a vague goal such as “make answers more accurate.” It hides several distinct obligations.

Also distinguish an assistant that explains a policy from a system that approves an expense or changes an employee record. For the hypothetical pilot below, keep the assistant read-only; that is a proposed product boundary, not a claim about any deployed system.

3. Establish the simplest prompting baseline

Before building retrieval, test a clear prompt with a small, curated evidence packet.

For a policy assistant, a proposed instruction set could be:

Answer only from the supplied approved policy excerpts. State the applicable scope and effective date. Cite the supporting passage. If evidence is missing or conflicts cannot be resolved using the stated precedence rules, explain the limitation and request review.

Add examples of good answers, missing-evidence responses, and unresolved conflicts. Reference text and examples are documented prompting techniques; they provide a baseline against which a more complex system can be compared. (developers.openai.com)

This baseline also exposes content problems. If a human reviewer cannot determine which policy applies, the next task is document governance—not model customization.

Do not assume that putting an entire document library into a large input solves the problem. Original research on long-context models found that performance could depend on where relevant information appeared, with weaker performance when it was placed in the middle. Test your own document lengths and layouts rather than adopting those results as a universal rule. (developers.openai.com)

Baseline deliverable: a versioned prompt, representative questions, expected evidence, and recorded outputs. Keep it available even if you later adopt RAG.

4. Choose RAG for changing, attributable knowledge

RAG separates the searchable knowledge collection from the model’s learned parameters. In a typical workflow, the application retrieves passages and includes them in the input used to generate an answer. (arxiv.org)

Prepare documents for that workflow, rather than simply uploading a folder and assuming the job is finished. For the proposed policy assistant, use this preparation checklist:

  • Identify approved documents and separate drafts.
  • Record effective dates, scope, versions, and superseded documents.
  • Preserve headings, exceptions, and table labels during extraction.
  • Attach stable identifiers so answers can point to the supporting material.
  • Define how updates enter the searchable collection.
  • Decide how access permissions will be enforced.

Some retrieval systems support metadata filtering before semantic search. OpenAI’s current retrieval documentation describes attribute filters, adjustable ranking, and hybrid search combining semantic and keyword matches. Those are useful implementation options—not substitutes for choosing authoritative documents. (developers.openai.com)

Check freshness operationally. A source-document update is not enough: confirm that ingestion completed and search returns the new passage. OpenAI also documents that file removal is eventually consistent, meaning removed content may remain searchable briefly. Test replacement and deletion behavior in whichever system you select. (developers.openai.com)

Finally, a citation is not proof of correctness. NIST recommends checking generated sources and citations during pre-deployment measurement and ongoing monitoring. A relevant-looking document can still fail to support the particular answer. (nvlpubs.nist.gov)

5. Consider fine-tuning for demonstrated behavior failures

Fine-tuning becomes a candidate when correct evidence is already available but the model repeatedly mishandles the task.

Documented supervised fine-tuning use cases include classification, specific output formats, and correcting instruction-following failures. It requires example inputs and known-good outputs, followed by evaluation against a baseline. (developers.openai.com)

For the proposed assistant, useful training demonstrations would show how to:

  • Preserve a qualifying exception.
  • Ask for missing location or employee-category information.
  • Explain an unresolved conflict without choosing arbitrarily.
  • Produce the required answer structure from supplied evidence.

Those are proposed demonstrations, not tested training results. The objective would be to teach an evidence-handling pattern—not to make changing policy facts authoritative merely because they appeared in training.

Keep training and evaluation examples separate. OpenAI’s fine-tuning guidance recommends splitting collected examples into training and test portions and retaining effective instructions in training examples. (developers.openai.com)

Availability check, October 4, 2026: OpenAI’s self-serve fine-tuning platform is no longer available to new users. Its published timeline says active existing customers will lose the ability to create new training jobs on January 6, 2027; inference continues until the underlying base model is deprecated. Fine-tuning remains a general technique, but this particular service is not a default option for a new project. (developers.openai.com)

Before committing elsewhere, verify the provider’s supported models, data terms, deployment options, and lifecycle.

6. Worked example: an internal-policy assistant

Hypothetical example—not a deployed product or measured result.

Assume a business wants employees to ask questions about an internal travel policy. The approved policy changes periodically, answers must cite evidence, and the assistant must not approve requests.

Create a fictional test collection containing:

  • An older policy marked as superseded.
  • A current policy stating that overnight travel requires advance manager approval.
  • A draft FAQ saying approval is unnecessary.
  • A reimbursement document that says nothing about contractors.

Under these assumptions, start with RAG and a clear evidence-handling prompt. Fresh documents and inspectable sources are the main requirements. Add fine-tuning only if the model continues to mishandle the retrieved evidence.

Use these test questions:

Test question Expected behavior under the fictional assumptions
“Do I need approval for overnight travel?” Retrieve the current approved policy and explain its advance-approval requirement.
“The old handbook says otherwise. Which applies?” Identify the superseded status and explain the documented precedence.
“Does this reimbursement rule cover contractors?” State that the supplied evidence does not establish contractor coverage.
“The FAQ disagrees with the policy. Can I skip approval?” Distinguish the draft FAQ from the approved policy.
“Two approved documents conflict and neither supersedes the other. What should I do?” Describe the conflict and request review rather than invent a hierarchy.

A passing answer to the first question could be:

Under the current approved travel policy, overnight travel requires advance manager approval. Evidence: Travel Policy, current version, “Approval requirements.” I can explain the rule, but I cannot approve the request.

Next, deliberately remove the current policy from the retrieved evidence. The desired response should change to a limitation, not remain confidently identical.

This exercise tests whether the assistant uses evidence rather than merely repeating a plausible policy answer. It does not establish production reliability.

7. Evaluate failures separately and budget for operations

Evaluate retrieval and answer generation separately. A document-based assistant can fail because the relevant passage never arrived, or because the model received it and answered incorrectly. Treat those as different repair jobs. (developers.openai.com)

For the proposed pilot, use a scorecard covering:

  • Evidence retrieval: was the expected passage supplied?
  • Answer support: are material claims supported?
  • Version handling: did the applicable document win?
  • Missing evidence: did the assistant acknowledge the gap?
  • Conflicts: did it follow documented precedence or escalate?
  • Usability: could the reader understand the answer and next step?

Use task-specific evaluation, human review, and repeated testing after changes. OpenAI’s evaluation guidance recommends these practices rather than relying on impressions that a system “seems to work.” Set thresholds for your task; do not copy illustrative vendor targets as universal standards. (developers.openai.com)

Keep security tests alongside quality tests. NIST identifies indirect prompt injection through retrieved data as a risk and recommends verifying that fine-tuning does not compromise safety controls. For this pilot, test documents containing instruction-like text and attempts to expose material outside the user’s permitted scope. These tests do not establish comprehensive security. (nvlpubs.nist.gov)

Budget for document preparation, retrieval storage, model calls, evaluation, monitoring, reviewer time, and updates—not just training or token charges. Treat this as a planning checklist; actual costs require your own workload and provider terms.

Before launch, confirm:

  • The baseline remains available for comparison.
  • Approved sources have owners and update procedures.
  • Missing and conflicting evidence produce acceptable responses.
  • Permissions are tested outside the prompt.
  • Changes trigger regression tests.
  • A named reviewer can handle unresolved questions.

Choose the simplest approach that passes your product-specific tests. Retrieval addresses access to evidence; fine-tuning targets demonstrated behavioral needs. Neither replaces a maintained source collection and a clear definition of an acceptable answer.