Implementation Guide

How to Build an Internal ChatGPT for Your Company: Three Routes and a Checklist

Staff already use ChatGPT, often on personal accounts. This guide compares three ways to give them an internal ChatGPT instead: buy business seats, assemble an open-source stack, or deploy a platform on your own infrastructure. It ends with a build checklist.

An internal ChatGPT for a company is a ChatGPT-style assistant that staff sign in to with their work account, that answers from company documents each person is allowed to read, and whose prompts and logs stay under company policy. You can buy it as business seats, assemble it from open-source parts on your own servers, or deploy a platform that bundles those parts inside your infrastructure.

What an internal ChatGPT needs to do

Staff mostly want what they already get from ChatGPT at home: quick drafts, summaries and answers. The company needs more than that. Before choosing a route, write down the requirements the result has to meet:

  • Company sign-in. People log in with their work identity, and leavers lose access on the day their account is disabled.
  • Answers from internal documents. Policies, contracts, manuals and tickets, with citations back to the source.
  • The asker’s permissions. Someone who cannot open a file in SharePoint or on a file share should not receive its contents in an answer.
  • Data handling you can defend. You know where prompts, files and logs are processed and stored, and for how long.
  • Admin control. IT decides which models, connectors and features are switched on, and for which groups.
  • A cost that scales sensibly. Seat plans grow with headcount; self-hosted capacity grows with usage.

The three routes below meet these requirements in different ways, with very different amounts of work.

Route 1: buy business seats

The fastest route is a vendor’s business plan. The two common choices are OpenAI’s ChatGPT Business or Enterprise, and Microsoft 365 Copilot for companies that already run on Microsoft 365.

ChatGPT Business is self-serve. A Standard seat lists at $20 per user per month billed annually or $25 billed monthly, with a 2-seat minimum, and new subscriptions are capped at 200 paid seats (list, verified October 2026). It includes SAML single sign-on, MFA and connectors to tools such as Microsoft 365, Google Drive, Slack and GitHub. OpenAI states that it does not train on Business workspace data by default.

ChatGPT Enterprise is sold through OpenAI’s sales team at custom prices. According to OpenAI’s plan comparison, it adds SCIM provisioning, role-based access controls, enterprise key management, admin-set retention and data residency in ten regions.

Microsoft 365 Copilot lists at $30 per user per month paid yearly and requires a qualifying Microsoft 365 plan (list, verified October 2026). Microsoft says Copilot only surfaces organisational data the user can at least view, and that prompts and responses are not used to train its foundation models.

What you give up is location. Prompts and retrieved documents are processed in the vendor’s cloud. Region settings decide where content is stored, and on some plans where models run, but the data still leaves your infrastructure. If some of your data cannot do that, this route covers only part of the workforce. Our explainer on whether ChatGPT is safe for confidential data goes through each plan’s data terms.

Route 2: assemble it yourself

The do-it-yourself route puts open-source components on servers you control. A typical stack has five layers:

LayerCommon choicesWhat it does
ModelOpen-weight models such as OpenAI’s gpt-oss, Llama, Qwen or MistralWrites the answers
Inference servervLLM or OllamaServes the model behind an OpenAI-compatible API
Chat interfaceOpen WebUIGives staff a ChatGPT-like web app
RetrievalAn embedding model, a vector store and connectors to your document sourcesFinds the passages each answer is grounded in
IdentityYour identity provider over OIDC or SAMLSigns people in and supplies their groups

A few published facts shape the choices. OpenAI releases its gpt-oss models under the Apache 2.0 licence, and the model card says gpt-oss-120b fits on a single 80 GB GPU while gpt-oss-20b runs within 16 GB of memory. vLLM’s server implements the Chat Completions, Completions, Embeddings and Responses APIs, so most chat front ends can talk to it unchanged. Ollama supports a subset of the OpenAI API and is the simpler choice for a single machine. If local LLMs are new to your team, start there.

Open WebUI describes itself as a self-hosted platform built to run entirely offline. Its feature list includes roles, groups and per-resource permissions, SSO/OIDC/LDAP sign-in and SCIM 2.0 provisioning. Read its licence before you rebrand it: since version 0.6.6 in April 2025, deployments with more than 50 users in a 30-day period must keep the Open WebUI branding unless they hold an enterprise licence.

A working demo takes little time. What remains is the work that makes the assistant safe for a whole company:

  • Permission-aware retrieval. Each document has to carry the access list from the SharePoint site, Confluence space or file share it came from, and every query has to be filtered by the asking user’s groups before ranking. Demos usually skip this, and security teams test it first. The guide to permission-aware retrieval covers the design.
  • Operations. GPU drivers, model upgrades, security patches, backups, monitoring and an on-call rota.
  • Evidence. Logs that show who asked what, which documents were retrieved and which model version answered.

This route gives full control and no per-seat fee. It suits a company with a platform team that is ready to run the service for years after the build.

Route 3: deploy a platform on your own infrastructure

The third route installs a product that bundles the same layers on servers you control: model serving, chat, retrieval, identity, audit and administration. It can run on-premises, in a private or sovereign cloud, or on an air-gapped network. Data stays where route 2 keeps it, and most of the integration work moves to the vendor.

Questions to put to any platform vendor:

  1. Does it run with no outbound internet connection at all?
  2. How does it import and enforce document permissions from each source system?
  3. Which identity providers does it support natively, and which need a proxy in front?
  4. Which models can it serve, and can you swap them without re-indexing your documents?
  5. What does a log record contain, and can you export it to your SIEM?
  6. Who patches which component, and how do upgrades reach an offline network?

The trade-off is a licence fee and a dependency on one vendor’s roadmap, in exchange for far less integration and maintenance work.

The three routes side by side

Buy seatsAssemble it yourselfPlatform on your infrastructure
Data controlProcessed in the vendor’s cloud; Enterprise plans add a storage regionOn your serversOn your servers, private cloud or air-gapped network
Identity and permissionsSAML SSO on ChatGPT Business and Enterprise; SCIM and RBAC on Enterprise; Copilot follows Microsoft 365 permissionsYou wire up SSO, groups and roles yourselfBuilt in; check which identity providers are native
Document accessVendor connectors to SaaS sourcesYou build connectors and permission syncBundled connectors; verify permission handling per source
Cost driversPer seat per month, plus credits for usage above included limitsGPUs, power and engineering timeLicence, GPUs, power and a lighter operations load
EffortLowest; ChatGPT Business is self-serveHighest; you own every layerModerate: install, identity and source onboarding

For the seat arithmetic at 50, 200 and 1,000 users, see what ChatGPT and Copilot seats cost per year.

Build checklist

Use this list whichever route you choose. Every item applies to bought seats too; the vendor simply answers some of them for you.

  1. Model choice. Shortlist two or three models. Check each licence for commercial use, then test them on 50 to 100 real questions collected from staff.
  2. GPUs. Size from model memory, context length and how many people ask at the same moment, not from total headcount. Weights set the floor; the KV cache for long, concurrent conversations sits on top.
  3. Retrieval with per-user permissions. Index sources together with their access lists, filter by the asker’s groups before ranking, and keep a test set of documents each department must never see. Private RAG explains the architecture.
  4. Identity. Single sign-on through your identity provider, groups mapped to roles, and same-day removal for leavers.
  5. Logging and audit. Record the user, the model and its version, the retrieved sources and the response. Set a retention period and send the logs to your SIEM.
  6. Acceptable-use policy. State which data classes may be entered, which tools are approved and what happens to work done in personal AI accounts.
  7. Pilot users. Start with two or three departments that have real document questions, such as HR policies, IT runbooks or product documentation. Collect the questions it fails every week.
  8. Adoption. Track weekly active users and answer quality. Announce a switch-off date for unapproved tools only once the internal assistant is part of daily work.

Which route fits which company

  • Buy seats when the data staff will use may go to a cloud vendor under contract, your documents already live in Microsoft 365 or mainstream SaaS tools, and speed matters more than control.
  • Assemble it yourself when you have a platform team, a narrow first use case and the appetite to run the service long term.
  • Deploy a platform on your infrastructure when data must stay inside your network, several departments need permission-aware answers, and auditors will ask for evidence. The private AI overview covers the deployment options.

The routes can also be combined: business seats for public and internal material, and a private assistant for the data that has to stay inside.

How VDF AI fits route 3

VDF AI Chat is a route 3 product: a ChatGPT-style assistant with private RAG that deploys on-premises, in a sovereign cloud region or on an air-gapped network. Connected sources such as SharePoint, Confluence and uploaded files bring their access lists into the index, and each query is filtered by the asking user’s identity. Every turn writes an audit record with the user, the model version and the sources retrieved.

It serves open-weight models inside your perimeter, plus any OpenAI-compatible endpoint you approve. On-premises, Microsoft Entra ID single sign-on is built in and maps Entra security groups to roles; Okta, Keycloak and other SAML or OIDC providers connect through an SSO-aware reverse proxy. Role-based access control is included on every plan.

For the VDF-specific setup, from connecting sources to citing answers, follow the step-by-step VDF AI walkthrough.

Sources

Frequently asked questions

What is an internal ChatGPT?

An internal ChatGPT is a chat assistant that a company provides to its own staff. People sign in with their work identity, the assistant answers from company documents they are allowed to read, and prompts, files and logs fall under company policy instead of personal accounts. It can be a vendor's business plan, an open-source stack you assemble on your own servers, or a platform deployed inside your own infrastructure.

Can a company build its own ChatGPT for internal use?

Yes. The usual self-built stack is an open-weight model, an inference server such as vLLM or Ollama, a chat interface such as Open WebUI, a retrieval pipeline over company documents, and single sign-on through your identity provider. A demo comes together quickly. The lasting work is permission-aware retrieval across every document source, security patching, model upgrades and keeping the service available for everyone who comes to depend on it.

How much does an internal ChatGPT cost?

It depends on the route. Seat plans are priced per user per month: at list price in October 2026, ChatGPT Business costs $20 per user per month billed annually and Microsoft 365 Copilot $30 paid yearly. A self-hosted assistant has no per-seat fee. Its cost comes from GPU servers, power, software licences and the people who run it, so it grows with usage capacity instead of headcount.

How do we stop an internal ChatGPT from showing people documents they should not see?

Enforce permissions at retrieval time. Each indexed document keeps the access list from its source system, and every query is filtered by the asking user's identity and group memberships before any passage reaches the model. Test it with accounts from different departments and a list of documents each one must never see. Filtering only in the chat interface, or trusting the system prompt to hold back, does not survive that test.

Should we block ChatGPT once our internal assistant is live?

Block it only after the internal assistant is in daily use and your acceptable-use policy says which data may go where. Blocking first leaves staff with personal phones and accounts as the only option, and you cannot see those at all. Once adoption is proven, restrict unapproved AI sites at the proxy, keep a short exception process for approved business plans, and tell people exactly which tool to use for which kind of data.

Filed under
private AIprivate RAGlocal LLMopen-weight modelson-premises AIenterprise AI
Private RAG & Search

Evaluate your knowledge stack

Find out how a private RAG and retrieval layer would perform on your data — accuracy, latency, governance, and what to fix before you scale.

Or start free — no credit card →

Keep reading