LOCAL LLM

Local LLM Build and
Implementation Support

What is a local LLM?A way to run generative AI like ChatGPT inside your own company, without relying on outside services. The data you enter never leaves your company.

Under your company’s control

APIA gateway that lets programs call the functions of an outside service. Cloud AI is used through an API.The line where cloud AI beginsBelow this line: vendors such as ChatGPT

Provided by
AI vendors
like ChatGPT

Operations & security

  1. People

    • Review and approval
    • Adoption
    We support
  2. UIThe screens people use to work with AI, such as chat or business app screens. (screens)

    • Chat
    • Business apps
    • Existing systems
  3. AgentAn AI that has a role for a specific task and gets the work done using set steps and tools.

    • Tasks and steps to delegate
    • Tools and permissions
  4. HarnessThe machinery that controls the model. It handles internal document search, tool connections, output checks and more.

    • Internal document search (RAGA mechanism that finds internal documents related to a question and has the AI answer based on their content.)
    • Tool connections (MCPA shared standard (connection method) for connecting AI to internal systems and tools.)
    • GuardrailsA mechanism that checks what goes into and comes out of the AI and stops errors or inappropriate content.
  5. ModelThe core of the AI, which understands and writes text. Also called a large language model (LLM).

    • Open modelsAn AI model whose contents are published and that can run on your own hardware. License terms differ by model.
    • QuantizationA technique that makes a model lighter so it runs with less memory. Answer quality may drop slightly.
    • Custom trainingGiving a model additional training on your own data. (if needed)
  6. Environment

    • GPU serversA server fitted with GPUs, the components that speed up AI computation.
    • Private cloudA cloud environment dedicated to your company and used cut off from the internet.
    • Inference engineSoftware that runs the model and produces answers. How much it can handle depends on its settings.
  7. Power

    • Power consumption estimates
    • Power supply and cooling
    • Power-saving settings
    We advise

Data

  • Internal documents
  • Databases
  • Business systems
SearchTrainingWe connect
With cloud AI such as ChatGPT, the AI vendor owns everything below the API. With a local LLM, every layer stays inside your company.

We handle every layer together, from choosing the hardware and model to the agents, the screens and adoption across your company. We also run LLMs in several environments of our own.

WHY LOCAL

Why Companies Consider a Local LLM

  1. “Our internal rules do not allow cloud AI”

  2. “Our contracts with business partners prohibit sending data outside”

  3. “Our work network is not connected to the internet”

  4. “We are not sure going fully local would be worth the cost”

Two office workers deep in thought over documents in a meeting

COMPARE

What Is a Local LLM: How It Differs from Cloud Generative AI

The difference is whether the data you enter leaves your company.The seven layers, explained with diagrams

Cloud generative AI

Data goes to an outside AI vendor

  • ProAccess to the latest models
  • ProQuick to start
  • CaveatConfidential data sometimes cannot be entered

Local LLM

Data stays inside your company

  • ProCan handle confidential data too
  • ProCosts rise little even as usage grows
  • CaveatMay fall short of the cloud on advanced reasoning
  • CaveatRequires setting up and running the environment

These are general tendencies. We confirm the difference for your own work with a parallel run.

HOW TO CHOOSE

How to Deploy a Local LLM: Three Setups Chosen per Task

You do not need to pick one setup for the whole company. You choose for each task.

  1. Cloud-first

    For work with no confidential data. We will not push you toward local.

  2. Hybrid

    Only the processing that involves confidential data runs in-house.

  3. Fully local (closed network)

    A setup that sends nothing outside at all.

When We Do Not Recommend a Local LLM

  • You do not handle confidential data, and cloud AI meets your internal rules
  • You need top-tier answer quality above all else
  • Your usage is low, so the cost is not justified

Find the Right Setup with 3 Questions

Answered 0 / 3

  1. Q1Do you handle information that must not leave the company (personal data, business partners’ confidential information, etc.)?
  2. Q2Do you need an environment cut off from the internet?
  3. Q3What do you want AI to handle?

LAYERS

The Seven Layers of a Local LLM and How We Support Them

They are listed in the same order as the diagram. Layers with no label are ones we build. Open a row to see the details and the evidence.

  1. PeoplePeople who check, and a setup that keeps it in use

    We support

    We decide where people check and where AI takes over, and make it stick with training and runbooks.

    • Designing review and approval
    • Training and handover

    From researchOn tasks AI is good at, speed rose by 25.1%; on tasks it is poor at, the rate of correct answers fell from 84.5% to 60–71%.[1]

  2. UI (screens)Screens where you can check the evidence

    We build screens that work inside your business systems and apps, where people can check the basis of an answer on the spot and correct it.

    • Embedding in business systems
    • Showing the sources

    From researchAdding citations has been shown to raise trust even when the cited sources have nothing to do with the content.[2]

  3. AgentDecide what to delegate, task by task

    We choose between a workflowA mechanism that runs AI and tools exactly according to set steps. that follows fixed steps and a setup where the AI picks the steps, based on the risk of the work.

    • Designing the tasks and steps to delegate
    • Setting the tools and permissions it may use

    From researchEven with state-of-the-art models, the share that succeeded every time when repeating the same task 8 times was under 25%.[3]

  4. HarnessControl the model with layer upon layer of checks

    On its own, an open modelAn AI model whose contents are published and that can run on your own hardware. License terms differ by model. only returns text. We control its output with RAGA mechanism that finds internal documents related to a question and has the AI answer based on their content., tool connections via MCPA shared standard (connection method) for connecting AI to internal systems and tools., and guardrailsA mechanism that checks what goes into and comes out of the AI and stops errors or inappropriate content..

    • Fixing and validating the output format
    • Checking inputs and tools

    From researchWith the same model, simply changing the design of the tools raised the task resolution rate from 11% to 18%.[4]

    Our track recordWe run our own internal agents on a local LLM, with read-only tools, limits on where they can be used, and activity logs built in.

  5. ModelCompare open modelsAn AI model whose contents are published and that can run on your own hardware. License terms differ by model. on your data and choose

    We compare and choose using questions from your own work, not rankings.

    • Checking Japanese-language quality and license terms
    • Checking quality after quantizationA technique that makes a model lighter so it runs with less memory. Answer quality may drop slightly.

    From researchIn a study that quantized a large model and compared it on harder questions, the drop in Japanese-language quality was 1.7% in automatic evaluation but 16.0% in human evaluation.[5]

  6. DataRAGA mechanism that finds internal documents related to a question and has the AI answer based on their content. first, before custom trainingGiving a model additional training on your own data.

    We connect

    We first check whether RAGA mechanism that finds internal documents related to a question and has the AI answer based on their content. is enough, and carry out additional training (such as LoRAA way to give a model additional training with little computation and hardware.) only when needed.

    • Connecting internal documents and databases
    • Respecting viewing permissions

    From researchIn comparisons of answering from company knowledge, RAGA mechanism that finds internal documents related to a question and has the AI answer based on their content. was consistently more accurate than training the model further on the documents as they are.[6]

  7. EnvironmentVendor-neutral hardware and inference engineSoftware that runs the model and produces answers. How much it can handle depends on its settings. selection

    We do not sell hardware, so we can choose without being tied to a manufacturer. We also run LLMs in several environments of our own.

    • Testing on private cloudA cloud environment dedicated to your company and used cut off from the internet. GPUsA component that speeds up AI computation. before you buy
    • Configuring the inference engineSoftware that runs the model and produces answers. How much it can handle depends on its settings.
    1. Try smallA few people take turnsA workstation with a GPUA component that speeds up AI computation.
    2. Department useDozens of people; internal document searchOne GPU serverA server fitted with GPUs, the components that speed up AI computation., or a private cloudA cloud environment dedicated to your company and used cut off from the internet.
    3. Company-wide useHundreds of people; response-time targetsMultiple GPU serversA server fitted with GPUs, the components that speed up AI computation., or a scalable private cloudA cloud environment dedicated to your company and used cut off from the internet.

    From researchOne report found that, on the same hardware, the design of the inference engineSoftware that runs the model and produces answers. How much it can handle depends on its settings. changed how much could be processed by 2 to 4 times.[7]

  8. PowerEstimate power use and cooling before the hardware

    We advise

    Before choosing hardware, we put power capacity, heat output and location into the requirements.

    • Estimating power consumption and cooling
    • Power-saving settings

    From researchOne report found that optimizing how models are run cut energy consumption by up to 73%.[8]

  9. Operations & securityEvaluation, monitoring, logging and permissions across every layer

    We switch models or settings only after comparing them in a parallel runGiving the same input to both the cloud and the local setup and comparing the results.. Everything is handled by our ISMSCertification under the international standard for information security management systems (ISO/IEC 27001).-certified team.

    • User permissions
    • Activity logs
    • Logs of data sent outside
    • Data flow diagram

    From researchIn one case, a cloud model with the same name saw its rate of correct answers on a task change from 84% to 51% within 3 months.[9]

Your Role and Ours

AreaWhat we ask of youWhat we take on
Work and dataSetting priorities and deciding what may be given to AIAn inventory and a proposed classification
Answer qualityJudging whether it is usable for the workWriting the test questions and tallying the results
Environment and powerDeciding the location and getting internal approvalHardware and power requirements, design and build
OperationsAn internal point of contactMonitoring, updates and handover

OUR PRACTICE

Our Own Example: Choosing Cloud or Local LLM Feature by Feature

In our own business tool, projectAI, we switch where each AI feature is processed. What we propose to you is this setup, which runs every day.

  1. Processing without confidential dataCloud onlyProcessed by cloud AI
  2. Processing being considered for migrationParallel runProcessed by both, and the results compared
  3. Processing with confidential dataLocal firstProcessed locally; sent to the cloud only when that is not possible, with the reason logged
Every request is loggedCost (yen)Cloud-to-local ratioReasons for sending to the cloud

Four Steps to Handing Work to Local

  1. Run it in the cloud
  2. Run local in parallel
  3. People compare the results
  4. Switch over the features that are ready to hand over

What We Run Today

  • Switching is per feature. It takes effect in about 30 seconds, without stopping the system
  • Every request sent to the cloud is logged with its reason
  • AI costs are recorded in yen by feature and by project

FLOW

How Local LLM Implementation Works

  1. About 1–2 weeks

    Sort the work and decide the setup

    We separate data that may leave the company from data that may not, and decide how AI is used for each task, the model and hardware setup, and the cost outlook.

    Even if you stop here, you keep the classification sheet and the proposed setup

  2. About 1 month or more

    Compare in a parallel run

    We give the same input to the cloud and to the local setup, and compare answer quality, speed and cost. You can test on a cloud GPU environment before buying hardware.

    The test questions can be reused for later model updates

  3. About 1–3 months

    Build the production environment

    We set up the environment and the agent features (internal document search, tools, verification, permissions, logging). If you buy hardware, procurement takes additional time.

    We also hand over a data flow diagram you can use for internal reviews

  4. Ongoing

    Operate it and hand it over in-house

    We take on monitoring and model updates. We prepare runbooks and hand over until your team can run it in-house.

    We can also continue running it for you

COST & DURATION

How We Think About Local LLM Cost and Duration

Local LLM costs basically range from several million yen to several billion yen. The cost varies greatly with the model used and the environment.

What Drives the Cost

It varies greatly with the model used and the environment

  • The model used (size and number)
  • The environment (hardware and location)
  • Number of users and concurrent users
  • Volume of documents handled
  • Integration with existing systems
  • Who runs operations

With a private cloud, you may be able to start without buying hardware.

Typical Duration

Through validation: from about 1.5 months

  • Sorting the work: 1–2 weeks
  • Validation in a parallel run: about 1 month or more
  • Building the production environment: about 1–3 months
  • Operations and handover: ongoing

If you buy hardware, procurement takes additional time.

Consult Us Free, Starting with How Much to Keep In-House

It is fine if you are not yet sure what your rules or contracts require.

Get a Free Consultation

FAQ

Frequently Asked Questions About Implementing a Local LLM

What is a local LLM?

It is a way of running generative AI on your own equipment or inside a cloud that your company manages. The information you enter is not sent to an outside AI vendor.

How is it different from using a cloud AI API?

With cloud AI such as ChatGPT, the AI vendor owns the layers below the APIA gateway that lets programs call the functions of an outside service. Cloud AI is used through an API. (from the agentAn AI that has a role for a specific task and gets the work done using set steps and tools. machinery down to the hardware and power), and your data is processed there. With a local LLM, every layer sits in an environment your company manages.

What can a local LLM do?

Searching internal documents and rules, drafting replies to inquiries, summarizing meeting minutes, classifying documents, and similar work. The more fixed the format of the work, the better it fits.

Does a local LLM remove the risk of information leaks?

It removes the risk of information going to an outside AI vendor. Internal permissions and activity logs still need to be designed separately, and we include them in what we build.

How does the quality of a local LLM’s answers compare with cloud generative AI?

For fixed-format work such as summarizing and search, it often reaches a sufficient level. It can fall short on advanced reasoning, so we confirm the gap with a parallel runGiving the same input to both the cloud and the local setup and comparing the results. before deciding.

Can we consult you from hardware selection onward?

Yes. The hardware you need depends on the size of the model, how many people use it at the same time, and the length of the documents it reads. With a private cloudA cloud environment dedicated to your company and used cut off from the internet. you can start without buying hardware, and you can also try it before you buy.

Can the model be trained on our own data?

Yes. We first check whether internal document search (RAGA mechanism that finds internal documents related to a question and has the AI answer based on their content.) is enough, and if it is not, we carry out additional training (such as LoRAA way to give a model additional training with little computation and hardware.).

Which model do you use?

We choose from open-source models based on Japanese-language quality, a size that fits your hardware, license terms (whether commercial use is allowed), and the ability to use tools. We decide after comparing them on your own work.

Can’t we just use an open-source model as it is?

You can if all you need is text back. To use it in your work, you need machinery for searching internal data, calling tools, checking output and so on. We handle its design and implementation as well.

How do we choose between on-premises and a private cloud?

If contracts or internal rules require storage on your own equipment, on-premisesPlacing and using hardware in your own buildings and facilities. is the right fit; if you want to start quickly and small, a private cloudA cloud environment dedicated to your company and used cut off from the internet. is the right fit.

How is the cost of implementing a local LLM determined?

It is basically from several million yen to several billion yen, and it varies greatly with the model used and the environment (hardware and where it is placed). We prepare an estimate after also sorting out the number of users and the scope of integration. Consultation is free.

Can it search internal documents and answer from them (RAG)?

Yes. It answers while showing the source documents, and it respects the viewing permissions of each document.

Can you also handle operations and help us bring it in-house?

We take on monitoring and model updates. We can also prepare runbooks and hand over until your team can run it in-house.

Last updated: