Engineering · AI

Public LLMs vs Fine-Tuned Private Models: Using AI Without Letting Your Company Data Out

Move fast with public AI APIs, or run isolated models on your own data? A comparison of the two approaches, the difference between RAG and fine-tuning, and a decision tree for choosing.

· ·8 minute read

The biggest dilemma in a company's AI transformation is this: integrate public AI services (such as OpenAI or Anthropic) via API and move fast, or use isolated, open-source models that work with your own data (such as Llama or Mistral)?

For a responsible manager or developer, sending company data, customer information or trade secrets to an AI service without safeguards carries serious regulatory (GDPR, KVKK) and intellectual-property risk. The best-known example is Samsung: after employees pasted confidential source code into a public chat tool, the company had to restrict the use of generative AI.

So how do you build enterprise AI without compromising data privacy? Let's look at the differences between public LLMs and customised models, and the criteria for choosing an architecture.

What are public and private LLMs?

A public LLM is a large language model that a provider runs on its own infrastructure and offers to everyone through an API. A private LLM runs on the company's own servers or in an isolated cloud environment, and is adapted to company data where needed.

Public LLM vs private model comparison

CriterionPublic cloud LLM (API)Private, isolated LLM (private AI)
Data privacyData is processed on the provider's infrastructureData stays on company servers or a private cloud
GDPR / KVKKRequires a data processing agreement (DPA) and a cross-border transfer reviewCompliance is easier because data stays under company control
Domain expertiseBroad general knowledge; can be shallow on company-specific workCan be adapted to company jargon and processes
Cost modelPay per use (tokens); grows with volumeFixed infrastructure / GPU cost; unit cost drops at high volume
Setup and maintenanceStart instantly; the provider maintains itSetup, updates and monitoring are the company's responsibility
Network dependencyResponse time depends on the external service and networkRuns on the internal network; no external dependency

2 core approaches to protect company data

Two main techniques are used in a customised AI architecture. Depending on the need, they can be combined.

1. RAG (Retrieval-Augmented Generation)

Instead of retraining the model, you connect a search layer built from your internal documents — usually a vector database — behind it.

  • How it works: when a user asks a question, the system first searches internal sources, finds the relevant information and passes it to the model with the instruction "answer only from these documents."
  • When to use it: the fastest and most economical option for FAQs, internal procedures, product catalogs and constantly changing data. When a document changes, so does the answer.

RAG works whether the model runs locally or behind a public API. The knowledge-base-grounded answers in oigodesk are a simple example of the approach.

2. Fine-tuning: teaching the model new skills

The parameters of an open-source language model are updated using thousands of high-quality examples your company has produced — past agent conversations, code repositories, expert reports and so on.

  • How it works: the model doesn't just learn facts; it adopts your company's tone, format and jargon.
  • When to use it: producing reports to a specific standard, writing code that fits your codebase, or matching industry jargon exactly.

RAG or fine-tuning?

For most companies the right starting point is RAG: it's cheaper, faster to set up and doesn't need retraining as your data changes. Fine-tuning comes in when you need to change how the model behaves rather than what it knows. Advanced solutions use both.

In the AI era, your edge isn't the model everyone can access — it's the data only you have.

Which should you choose? A decision tree

Question 1: Is there sensitive customer data or trade secrets?

  • Yes: don't send this data to public APIs without additional safeguards. Either mask and anonymise personal data and settle the data processing agreement and region, or run open-source models such as Llama, Mistral or DeepSeek on your own servers (on-premise) or in a private cloud (VPC) with a RAG architecture.
  • No: for general summarisation or marketing content, OpenAI or Anthropic APIs offer speed and cost advantages.

Question 2: How much volume do you have?

  • Low or medium volume: token-based public APIs are usually more economical, with no infrastructure or maintenance burden.
  • Millions of tokens a month: token costs grow quickly. Running a private model on your own GPUs or rented cloud GPUs can lower unit cost significantly. Compare on total cost, including GPUs, maintenance and the team.

The middle path: a hybrid architecture

For many companies the most practical answer is a mix: steps that handle sensitive data run on a local model, while general, non-sensitive work runs on a public API. A data-masking layer in between keeps personal information from leaving.

Frequently asked questions

What is RAG?

RAG (Retrieval-Augmented Generation) is a method where a language model searches the company's own documents before answering and bases its reply only on what it finds. It doesn't require retraining the model.

What is fine-tuning?

Fine-tuning is retraining an existing language model on company-specific examples so it picks up the company's tone, format and jargon.

Do public AI APIs use my data to train their models?

Major providers state that their business APIs don't use data sent through the API for model training by default. The data is still processed on the provider's infrastructure, though, so check the contract terms, retention periods and the region where data is processed.

What should I watch for with AI and data protection law?

If personal data is sent to a service in another country, cross-border transfer rules (under GDPR, KVKK and similar laws) apply. Masking personal data, putting the required agreements in place, or running sensitive steps on a local model reduces the risk. Consult your legal advisor for a definitive assessment.

Conclusion: your company's most valuable asset is its data

In the AI era, the real competitive advantage isn't the public APIs your competitors can also use; it's the unique data you've built up over the years. Building a secure AI architecture that processes this data without it leaving your control is a strategic necessity.

At oigoworks we build AI integrations that fit companies' data privacy requirements: assistants that answer from your own documents with RAG, models running locally or in a private cloud, and fine-tuning on company data where needed. For choosing the right approach, see also our guide custom or off-the-shelf software?

Keep your data yours.

Let's work out together which data should be processed where, and set up RAG, a local model or a hybrid architecture to fit your business.