BLOG
What Is a Local LLM? How It Differs from ChatGPT, Explained with Diagrams

A local LLM is a way of running the core of generative AI (the model) on equipment your company manages. You can use it without sending your questions or internal documents to an outside AI vendor.

This article uses diagrams to explain how a local LLM differs from ChatGPT and the seven layers that support it.
What Is a Local LLM?
An LLM (large language model) is the core of an AI that understands and writes text. An LLM is also what runs inside ChatGPT.
"Local" means a place your company manages. It can be any of these three:
- A single PC
- Servers inside your company (on-premises)
- A cloud environment dedicated to your company (private cloud)
How It Differs from ChatGPT: Who Owns How Much
Generative AI is made up of seven layers. From the top, they are people, UI (screens), agent, harness, model, environment and power.

When you use cloud AI such as ChatGPT from a business system, the AI vendor owns everything below the API (the gateway for calling an outside service). With a local LLM, every layer stays inside your company.
Note that many vendors of cloud AI used through an API do not use your input for training by default[1]. The difference lies in where the data is processed and who manages it.
What Each of the Seven Layers Does
How AI works is easier to understand if you compare it to a workplace. The model is a knowledgeable new hire, the harness is how the workplace runs, and people are the approvers.

Unlike a new hire, however, AI sometimes gives a plausible-sounding answer when it does not know. It also does not remember anything from one conversation to the next. That is why you need people who check and a harness.
A Model Alone Is Not Enough for Business Use
An openly available model, used as it is, only returns text. It does not know your internal documents, cannot use tools, and does not check its own answers.

The harness makes up for this. A harness was originally a piece of horse tack: gear that carries the horse's pulling power in the intended direction and holds it back from heading somewhere dangerous. One study found that, with the same model, simply changing the design of the tools raised the task resolution rate from 11% to 18%[2].
Pros and Cons
Pros
- You can use it without sending data to an outside AI vendor
- Costs rise little even as usage grows
- You can choose the model and setup to fit your own work
Cons
- Hardware and electricity cost money
- Your company has to handle operations and updates itself
- It may fall short of the latest cloud AI on advanced reasoning
Going local does not by itself make things secure. Internal viewing permissions and activity logs still need to be designed separately.
What Can a Local LLM Do?
- Search internal documents and rules, and answer from them
- Summarize meeting minutes and reports
- Draft replies to inquiries
- Classify and check documents
The more fixed the format of the work, the better it fits.
Choose the Right Setup for Your Company with Two Questions

You do not need to pick one setup for the whole company. You choose for each task. On our Local LLM Build and Implementation Support page, you can find the right setup with three questions.
Frequently Asked Questions
What is the difference between a local LLM and an LLM?
An LLM is the AI itself, which understands and writes text. A local LLM refers to running that LLM in a place your company manages.
Can I run ChatGPT locally?
You cannot run ChatGPT itself. However, many models have been released so that they can run on your own hardware, and OpenAI also released one in August 2025[3].
How well does a local LLM perform?
For fixed-format work such as summarizing and search, it often reaches a sufficient level. Small models that run on modest hardware can fall short of the latest cloud AI, so compare them on the same task before deciding.
Can a local LLM be used without the internet?
Once you have downloaded the model, it can run on your own hardware alone. You need a connection to get or update the model.
How much does a local LLM cost in electricity?
It is determined by power consumption × hours of use × the electricity rate. For example, running a top-of-the-line consumer GPU (up to 575 W) at full load for 8 hours a day comes to about ¥140 a day at a reference rate of ¥31/kWh[4][5].
Are there local LLMs that are strong in Japanese?
Yes. Some models are released with a focus on Japanese, such as LLM-jp from the National Institute of Informatics[6]. Making a model lighter to run can lower the quality of its Japanese, so test it with real questions from your work before choosing.
Are local LLMs free?
Many models are free to use, but hardware, electricity and operations still cost money. Whether commercial use is allowed differs by model, so check the license terms as well.
Summary
A local LLM is a way of keeping all seven layers of generative AI inside your company. Your data does not have to leave, but your company has more to prepare itself, from the hardware to how people use it.
We also switch between cloud and local for each AI feature of our own business tool. How it works, and what we learned from running it, is in our in-house case study. You can consult us on all seven layers together, from power and hardware requirements to the harness, the screens and adoption across your company. For details, see Local LLM Build and Implementation Support.
References
- OpenAI, "Data controls in the OpenAI platform"
- J. Yang et al., "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering," NeurIPS 2024
- OpenAI, "Introducing gpt-oss," August 2025
- NVIDIA, "Compare Current and Previous GeForce Series of Graphics Cards"
- Home Electric Appliances Fair Trade Conference, "FAQ" (in Japanese; reference electricity rate. This is a guideline for households and differs from corporate contract rates)
- National Institute of Informatics, "LLM-jp"