Home
Home
German Version
Support
Impressum
26.4 Release ►

Start Chat with Collection

    Main Navigation

    • Preparation
      • Connectors
      • Create an InSpire VM on Hyper-V
      • Initial Startup for G7 appliances
      • Setup InSpire G7 primary and Standby Appliances
    • Datasources
      • Configuration - Atlassian Confluence Connector
      • Configuration - Atlassian Confluence REST Connector
      • Configuration - Best Bets Connector
      • Configuration - Box Connector
      • Configuration - COYO Connector
      • Configuration - Data Integration Connector
      • Configuration - Database Connector
      • Configuration - Documentum Connector
      • Configuration - Dropbox Connector
      • Configuration - Egnyte Connector
      • Configuration - GitHub Connector
      • Configuration - Google Drive Connector
      • Configuration - GSA Adapter Service
      • Configuration - HL7 Connector
      • Configuration - IBM Lotus Connector
      • Configuration - Jira Connector
      • Configuration - JVM Launcher Service
      • Configuration - LDAP Connector
      • Configuration - Microsoft Azure Principal Resolution Service
      • Configuration - Microsoft Dynamics CRM Connector
      • Configuration - Microsoft Exchange Connector
      • Configuration - Microsoft File Connector (Legacy)
      • Configuration - Microsoft File Connector
      • Configuration - Microsoft Graph Connector
      • Configuration - Microsoft Loop Connector
      • Configuration - Microsoft Project Connector
      • Configuration - Microsoft SharePoint Connector
      • Configuration - Microsoft SharePoint Online Connector
      • Configuration - Microsoft Stream Connector
      • Configuration - Microsoft Teams Connector
      • Configuration - Salesforce Connector
      • Configuration - SCIM Principal Resolution Service
      • Configuration - SemanticWeb Connector
      • Configuration - ServiceNow Connector
      • Configuration - Web Connector
      • Configuration - Yammer Connector
      • Data Integration Guide with SQL Database by Example
      • Indexing user-specific properties (Documentum)
      • Installation & Configuration - Atlassian Confluence Sitemap Generator Add-On
      • Installation & Configuration - Caching Principal Resolution Service
      • Installation & Configuration - Mindbreeze InSpire Insight Apps in Microsoft SharePoint On-Prem
      • Mindbreeze InSpire Insight Apps in Microsoft SharePoint Online
      • Mindbreeze Web Parts for Microsoft SharePoint
      • User Defined Properties (SharePoint 2013 Connector)
      • Whitepaper - Migration of Sites Selected Permissions for the MS SharePoint Online Connector
      • Whitepaper - Migration of Tenant-Wide Permissions for the MS SharePoint Online Connector
      • Whitepaper - Mindbreeze InSpire Insight Apps in Salesforce
      • Whitepaper - Overview of Connectors
      • Whitepaper - Web Connector - Setting Up Advanced Javascript Usecases
    • Configuration
      • CAS_Authentication
      • Configuration - Advanced Configuration for Mail Delivery
      • Configuration - Alerts
      • Configuration - Alternative Search Suggestions and Automatic Search Expansion
      • Configuration - Back-End Credentials
      • Configuration - Chinese Tokenization Plugin (Jieba)
      • Configuration - CJK Tokenizer Plugin
      • Configuration - Collected Results
      • Configuration - CSV Metadata Mapping Item Transformation Service
      • Configuration - Entity Recognition
      • Configuration - Exporting Results
      • Configuration - Filter Plugins
      • Configuration - GSA Late Binding Authentication
      • Configuration - Identity Conversion Service - Replacement Conversion
      • Configuration - InceptionImageFilter
      • Configuration - Index-Servlets
      • Configuration - InSpire AI Chat and Insight Services for Retrieval Augmented Generation
      • Configuration - Item Property Generator
      • Configuration - Japanese Language Tokenizer
      • Configuration - JavaScript Transformer Plugins
      • Configuration - Kerberos Authentication
      • Configuration - Management Center Menu
      • Configuration - Metadata Enrichment
      • Configuration - Metadata Reference Builder Plugin
      • Configuration - Mindbreeze Proxy Environment (Remote Connector)
      • Configuration - Personalized Relevance
      • Configuration - Plugin Installation
      • Configuration - Principal Validation Plugin
      • Configuration - Profile
      • Configuration - Reporting Query Logs
      • Configuration - Reporting Query Performance Tests
      • Configuration - Request Header Session Authentication
      • Configuration - Shared Configuration (Windows)
      • Configuration - Vocabularies for Synonyms and Suggest
      • Configuration of Thumbnail Images
      • Cookie-Authentication
      • Documentation - Mindbreeze InSpire
      • I18n Item Transformation
      • Installation & Configuration - Outlook Add-In
      • Installation - GSA Base Configuration Package
      • JWT Authentication
      • Language detection - LanguageDetector Plugin
      • Mindbreeze Personalization
      • Mindbreeze Property Expression Language
      • Mindbreeze Query Expression Transformation
      • SAML-based Authentication
      • Trusted Peer Authentication for Mindbreeze InSpire
      • Using the InSpire Snapshot for Development in a CI_CD Scenario
      • Whitepaper - AI Chat
      • Whitepaper - Create a Google Compute Cloud Virtual Machine InSpire Appliance
      • Whitepaper - Create a Microsoft Azure Virtual Machine InSpire Appliance
      • Whitepaper - Create AWS 10M InSpire Appliance
      • Whitepaper - Create AWS 1M InSpire Appliance
      • Whitepaper - Create AWS 2M InSpire Appliance
      • Whitepaper - Create Oracle Cloud 10M InSpire Application
      • Whitepaper - Create Oracle Cloud 1M InSpire Application
      • Whitepaper - MMC_ Services
      • Whitepaper - Single Sign-On with Microsoft Entra ID or Active Directory Federation Services
      • Whitepaper - Text Classification Insight Services
    • Operations
      • Adjusting the InSpire Host OpenSSH Settings - Set LoginGraceTime to 0 (Mitigation for CVE-2024-6387)
      • app.telemetry Statistics Regarding Search Queries
      • Blacklisting vulnerable kernel modules esp4, esp6, rxrpc - (Mitigation for CVE-2026-43284 _ DirtyFrag)
      • CIS Level 2 Hardening - Setting SELinux to Enforcing mode
      • Configuration - app.telemetry dashboards for usage analysis
      • Configuration - Usage Analysis
      • Disabling algif_aead_init - (Mitigation for CVE-2026-31431)
      • FAQ - Creating Mindbreeze InSpire Appliances on Hyper Scalers
      • Handbook - Backup & Restore
      • Handbook - Command Line Tools
      • Handbook - Distributed Operation (G7)
      • Handbook - Filemanager
      • Handbook - Indexing and Search Logs
      • Handbook - Updates and Downgrades
      • Index Operating Concepts
      • Inspire Diagnostics and Resource Monitoring
      • Provision of app.telemetry Information on G7 Appliances via SNMPv3
      • Restoring to As-Delivered Condition
      • Whitepaper - Administration of Insight Services for Retrieval Augmented Generation
      • Whitepaper - Insight Workplace
      • Whitepaper - Mindbreeze in Microsoft Teams
      • Whitepaper - Mindbreeze in OpenAI ChatGPT
      • Whitepaper - Mindbreeze InSpire LLM_ A Kubernetes Integration Guide
      • Whitepaper - Mindbreeze InSpire LLM_ On-Premise Deployment Guide
      • Whitepaper - Natural Language Question Answering (NLQA)
      • Whitepaper - Overview AI based Document Parsing, Transcriptions, and Semantic Index Pipeline
      • Whitepaper - Overview of Agentic AI and the Insight Workplace
      • Whitepaper - Using the Mindbreeze InSpire MCP Server
    • User Manual
      • Browser Extension
      • Cheat Sheet
      • iOS App
      • Keyboard Operation
    • SDK
      • api.chat.v1beta.generate Interface Description
      • api.v2.alertstrigger Interface Description
      • api.v2.export Interface Description
      • api.v2.personalization Interface Description
      • api.v2.search Interface Description
      • api.v2.suggest Interface Description
      • api.v3.admin.SnapshotService Interface Description
      • Debugging (Eclipse)
      • Developing an API V2 search request response transformer
      • Developing Item Transformation and Post Filter Plugins with the Mindbreeze SDK
      • Developing Item Transformation Launched Service with Mindbreeze SDK
      • Development of a Query Expression Transformer
      • Development of Insight Apps
      • Embedding the Insight App Designer
      • Export and Integration of Personalization and Analytics Data with External Platforms
      • Java API Interface Description
      • OpenAPI Interface Description
      • SDK Overview
    • Release Notes
      • Release Notes 20.1 Release - Mindbreeze InSpire
      • Release Notes 20.2 Release - Mindbreeze InSpire
      • Release Notes 20.3 Release - Mindbreeze InSpire
      • Release Notes 20.4 Release - Mindbreeze InSpire
      • Release Notes 20.5 Release - Mindbreeze InSpire
      • Release Notes 21.1 Release - Mindbreeze InSpire
      • Release Notes 21.2 Release - Mindbreeze InSpire
      • Release Notes 21.3 Release - Mindbreeze InSpire
      • Release Notes 22.1 Release - Mindbreeze InSpire
      • Release Notes 22.2 Release - Mindbreeze InSpire
      • Release Notes 22.3 Release - Mindbreeze InSpire
      • Release Notes 23.1 Release - Mindbreeze InSpire
      • Release Notes 23.2 Release - Mindbreeze InSpire
      • Release Notes 23.3 Release - Mindbreeze InSpire
      • Release Notes 23.4 Release - Mindbreeze InSpire
      • Release Notes 23.5 Release - Mindbreeze InSpire
      • Release Notes 23.6 Release - Mindbreeze InSpire
      • Release Notes 23.7 Release - Mindbreeze InSpire
      • Release Notes 24.1 Release - Mindbreeze InSpire
      • Release Notes 24.2 Release - Mindbreeze InSpire
      • Release Notes 24.3 Release - Mindbreeze InSpire
      • Release Notes 24.4 Release - Mindbreeze InSpire
      • Release Notes 24.5 Release - Mindbreeze InSpire
      • Release Notes 24.6 Release - Mindbreeze InSpire
      • Release Notes 24.7 Release - Mindbreeze InSpire
      • Release Notes 24.8 Release - Mindbreeze InSpire
      • Release Notes 25.1 Release - Mindbreeze InSpire
      • Release Notes 25.2 Release - Mindbreeze InSpire
      • Release Notes 25.3 Release - Mindbreeze InSpire
      • Release Notes 25.4 Release - Mindbreeze InSpire
      • Release Notes 25.5 Release - Mindbreeze InSpire
      • Release Notes 25.6 Release - Mindbreeze InSpire
      • Release Notes 25.7 Release - Mindbreeze InSpire
      • Release Notes 25.8 Release - Mindbreeze InSpire
      • Release Notes 26.1 Release - Mindbreeze InSpire
      • Release Notes 26.2 Release - Mindbreeze InSpire
      • Release Notes 26.3 Release - Mindbreeze InSpire
      • Release Notes 26.4 Release - Mindbreeze InSpire
    • Security
      • Known Vulnerablities
    • Product Information
      • Product Information - Mindbreeze InSpire - Standby
      • Product Information - Mindbreeze InSpire
    Home

    Path

    Sure, you can handle it. But should you?
    Let our experts manage the tech maintenance while you focus on your business.
    See Consulting Packages

    Whitepaper
    Mindbreeze InSpire LLM: On-Premise Deployment Guide

    IntroductionPermanent link for this heading

    Mindbreeze InSpire LLM lets you run large language models directly on a GPU-enabled InSpire appliance, entirely within your own infrastructure. Rather than routing requests to an external, hosted provider, models are deployed, served, and consumed on-premise – keeping prompts, source documents, and generated output inside your own security and network boundary, and giving you full control over which models run and who may use them.

    On-premise deployments are managed through the InSpire LLM CLI (inspire-llm), which drives the full lifecycle of a model on the appliance. The CLI is built around three concepts:

    • Template – a base configuration for deploying a particular model.
    • Provider – a parameterized template that fixes the GPU(s) device(s), the maximum token limit, whether the model is published (production) or unpublished (testing/development), and custom deployment arguments (e.g., GPU utilization).
    • Container – the Docker container that actually runs the model built from a chosen provider.

    Once started, a model is exposed over authenticated HTTPS endpoints on the appliance and can be consumed directly or through the InSpire RAG service. Access is secured with Keycloak using OAuth 2.0 / JWT tokens and role-based authorization, so only holders of the appropriate service-account roles – InSpire LLM Service User for published models, InSpire LLM Unpublished Services User for unpublished ones – can reach a given model.

    This guide walks through everything required to stand up such a deployment: preparing the GPU host and NVIDIA driver container, configuring the Keycloak service account, installing the InSpire LLM package, and creating, starting, accessing, and cleaning up model deployments — with CLI reference, health checks, and troubleshooting in the appendix.

    PrerequisitesPermanent link for this heading

    The InSpire LLM container can be deployed on a GPU-enabled Hardware, where the NVIDIA driver container is up to date.

    Installation process of the NVIDIA driver containerPermanent link for this heading

    The correct NVIDIA driver container image matching the kernel version is required:

    1. Install the image before updating to an InSpire version with a new kernel.

    docker load -it /path/to-driver-image.tar

    1. Install the update and the driver will be loaded automatically at boot.

    If the driver for the current kernel is not installed, the driver can also be loaded manually (after loading the image) to be available without reboot.

    systemctl start nvidia-driver

    Setup of Service AccountPermanent link for this heading

    The HTTPD configuration allows only users with the role “InSpire LLM Services User” and “InSpire LLM Unpublished Services User” using OAuth2.0/JWT-tokens to access the LLMs. The difference between those roles is:

    • “InSpire LLM Services User” can only use published model providers.
    • “InSpire LLM Unpublished Services User” may also use unpublished models.

    Additionally, the user name has to be part of the URL. For example:

    https://<appliance-fqdn>:8443/api/llm/<customer-user-name>/<llm-name>

    Currently, the Chat Service uses the service account of the client, therefore a client has to be created for each customer instead of a user.

    Prepare Realm RolePermanent link for this heading

    In the Mindbreeze InSpire MMC Console navigate to “Setup” and then “Credentials”.

    Please create the realm role in the Keycloak realm where the service accounts will be created (default realm is “master”):

    • “InSpire LLM Services User”

    Add OIDC Client for the Service AccountPermanent link for this heading

    Please add the client in the “Clients” section:

    Setup of Client ParametersPermanent link for this heading

    The OIDC Client should be enabled for “Client Authentication” and support the “Service Accounts roles” Authentication Flow.

    Both settings can be configured during client creation as shown in the screenshot or just be set after the client is configured.

    Assigning Service Account RolePermanent link for this heading

    On the “Service Account Roles” tab of the OIDC client settings, add the previously created “InSpire RAG Impersonating User” role to the service account of the client.

    Exporting Realm Access Roles in the OAuth Access TokensPermanent link for this heading

    For checking the required role for impersonation, the Mindbreeze InSpire RAG Service is expecting the realm roles as “roles” attribute of the OAuth access token.

    This requires the following configuration steps on the previously created OIDC client:

    1. Navigate to the “Client scopes” tab and open the settings of the scope marked as “Dedicated scope and mappers for this client”.

    1. Add the “realm roles” mapper from the “Predefined Mappers” by clicking on “Add Mapper” and selecting the “From Predefined Mappers” mapper type.

    1. Edit the added mapper, and change the Token Claim Name attribute to “roles” so that the service user realm roles are written as “roles” claim in the access token.

    1. Save your changes

    Installation ProcessPermanent link for this heading

    PrerequisitesPermanent link for this heading

    Before installing the InSpire LLM Installer (version 25.6), it is recommended to remove any previous InSpire LLM installations to ensure a clean environment. Follow these steps to remove existing components:

    1. Remove system service files:

    sudo rm /etc/systemd/system/llm-proxy.service

    sudo rm /etc/systemd/system/llm-proxy-test.service

    1. Stop and remove all running InSpire LLM Docker containers:

    docker stop <inspire-llm-container-name>

    docker rm <inspire-llm-container-name>

    1. Delete the InSpire LLM installation directory (default path: /var/data/inspire-llm):

    sudo rm -rf /var/data/inspire-llm

    InstallationPermanent link for this heading

    To install InSpire LLM, follow the steps below:

    1. Copy the installation archive to the target host, typically under /var/data.
    2. Set the installer as executable:

    chmod +x /var/data/inspire-llm-<version>-<hash-value>.sh

    1. Execute the installer:

    cd /var/data

    ./inspire-llm-<version>-<hash-value>.sh

    Note:

    During installation, all packaged Docker images and LLM models will be extracted to /var/data/inspire-llm by default. If you want to change the destination directory, execute the installer the following way:

    ./inspire-llm-<version>-<hash-value>.sh --dest-dir <destination-directory>

    (Optional) Make the InSpire CLI Globally AvailablePermanent link for this heading

    By default, the InSpire LLM CLI script is located at:

    /var/data/inspire-llm/bin/inspire-llm.sh

    To avoid specifying the full path every time you use the CLI, you can create a symbolic link in /usr/local/bin. This will allow you to run the CLI from any directory.

    To the create the symbolic link:

    sudo ln -s /var/data/inspire-llm/bin/inspire-llm.sh /usr/local/bin/inspire-llm

    After creating the symbolic link, you can use the InSpire LLM CLI simply by running:

    inspire-llm

    from any location in your terminal.

    Deploying a Large Language ModelPermanent link for this heading

    Creating a ProviderPermanent link for this heading

    InSpire LLM uses the concepts of templates and providers:

    • A template is a base configuration used to deploy a large language model.
    • A provider is a parameterized template that specified:
      • The GPU device(s) for deployment
      • The maximum allowed tokens for the model
      • Whether the model is published or unpublished

    You can create a provider either through the CLI (with options) or interactively:

    1. List available templates:

    inspire-llm list

    Copy the desired template name from the output

    1. Create a provider interactively:

    inspire-llm create-provider -t <template-name>

    Or,

    Create a provider non-interactively:

    inspire-llm create-provider -t <template-name> -g <gpu-device> --max-tokens <max-tokens> --published | --unpublished --extra-arguments “<extra-arguments>”

    1. Verify provider creation:

    inspire-llm list

    Your new provider should appear in the list.

    Deploying a Single Model on Multiple GPUsPermanent link for this heading

    To deploy a single model across multiple GPUs, specify the -g option with a comma-separated list of GPU device IDs. For example:

    inspire-llm create-provider -t <template-name> -g 0,1 –-max-tokens 24576 --published

    Published vs. Unpublished ProvidersPermanent link for this heading

    • A published provider exposes the model at the endpoint:

    http://llm-proxy/api/llm/...

    • An unpublished provider exposes the model at:

    http://llm-proxy/api/llm-instance/...

    Use published for production deployments and unpublished for testing and development environments.

    Creating a Container for the LLMPermanent link for this heading

    Once you have created a provider, you can create a container to run the LLM.

    1. List available providers:

    inspire-llm list

    Copy the provider name you want to use.

    1. Create a container:

    inspire-llm create-container -p <provider-name>

    This will create a Docker container for the specified provider.

    Starting a Large Language ModelPermanent link for this heading

    After creating a container, you can start it:

    1. List available containers:

    inspire-llm available-containers

    Copy the container name.

    1. Start the container:

    inspire-llm start-container -c <container-name>

    This will start the LLM container and refresh proxy configurations.

    1. Verify deployment:

    Check the container logs to ensure the server started successfully:

    docker logs -f <container-name>

    Note: Starting a large language model may take several minutes.

    Cleaning Up DeploymentsPermanent link for this heading

    InSpire LLM provides commands to stop and delete deployments when needed.

    1. List running containers:

    inspire-llm running-containers

    Copy the name of the container you wish to stop or delete.

    1. Stop a container:

    inspire-llm stop-container -c <container-name>

    1. Delete a container:

    inspire-llm delete-container -c <container-name>

    If the container is running, it will be stopped, the proxy configuration will be refreshed, and then the container will be deleted.

    1. Delete a provider (and its containers):

    inspire-llm delete-container -c <container-name>

    This will stop and delete any associated containers, refresh the proxy configuration, and remove the provider.

    Note:

    • Always ensure you use the correct names for templates, providers, and containers as listed by the inspire-llm list, inspire-llm available-containers and inspire-llm running-containers commands.
    • Published models should be used for production endpoints; unpublished for development and testing.

    Accessing a Deployed ModelPermanent link for this heading

    After deploying a model, it is accessible via HTTP endpoints:

    • If the InSpire container is running on your local machine:

    https://<appliance-fqdn>:8443/api/llm/<customer-user-name>/meta-llama-3.1-8b-instruct

    Or, if unpublished:

    https://<appliance-fqdn>:8443/api/llm-instance/<customer-user-name>/meta-llama-3.1-8b-instruct

    (Replace the <appliance-fqdn> with the appliance hostname you have deployed a model on)

    • If you are accessing through the proxy on the machine:

    http://llm-proxy:8000/api/llm/meta-llama-3.1-8b-instruct

    Or, if unpublished:

    http://llm-proxy:8000/api/llm-instance/meta-llama-3.1-8b-instruct

    Note:

    • <customer-user-name> may be any value.

    AppendixPermanent link for this heading

    InSpire LLM CLI CommandsPermanent link for this heading

    Command

    Explanation

    inspire-llm start-container -c <container-id>

    Start a container and refresh the proxy configuration.

    inspire-llm stop-container -c <container-id>

    Stop a container and refresh the proxy configuration.

    inspire-llm create-container -p <provider-name>

    Create a container from the specified provider.

    inspire-llm delete-container -c <container-id>

    Delete a container. If running, the container will be stopped and deleted. Also refreshes the proxy configuration.

    inspire-llm create-provider -t <template> -g <gpu-device> --max-tokens <max-tokens> --published | --unpublished –-extra-arguments “<extra-arguments>”

    Create a provider from the specified template. If only the template is specified, an interactive questionnaire is shown to set other parameters.

    inspire-llm delete-provider -p <provider-name>

    Delete a provider. If the associated container is running, it will be stopped and deleted, the proxy configuration will be refreshed, and the provider will be removed.

    inspire-llm available-containers

    List all available InSpire LLM containers (created or running models).

    inspire-llm running-containers

    List all currently running InSpire LLM containers.

    inspire-llm initproxy

    Initialize the proxy service.

    inspire-llm reloadproxy

    Restart the proxy service

    inspire-llm list

    List available templates and providers.

    inspire-llm load-docker-images

    Load Docker images required by InSpire LLM. Note: This command is executed during installation and usually does not need to be run manually.

    inspire-llm create-templates

    Create templates from available models. Note: This command is executed during installation and usually does not need to be run manually.

    inspire-llm copy-llm-proxy-service

    Copy llm-proxy.service to /etc/systemd/system. Note: This command is executed during installation and usually does not need to be run manually.

    inspire-llm copy-httpd-to-inspire

    Copy the HTTPD service to the InSpire container. Note: This command is executed during installation and usually does not need to be run manually.

    inspire-llm update-config

    Automatically update the config. Config contains the tags of the required docker images to deploy LLM.

    inspire-llm --help

    Display help and usage information.

    LLM Deployment Health ChecksPermanent link for this heading

    You can verify the health and status of your LLM deployments using one of the following endpoints. The correct URL depends on how your deployment is exposed.

    InSpire HostnamePermanent link for this heading

    If your LLM deployment is exposed through a valid InSpire container, use the following URL:

    https://<inspire-hostname>/api/llmhealth/<model-name>/health

    Direct host or IP addressPermanent link for this heading

    If you did not use InSpire to expose your deployment and are connecting directly to the host or IP address, use this URL:

    http://<host-or-ip-address>:<port>/api/llm/<model-name>/health

    Reverse ProxyPermanent link for this heading

    If you are checking the health from inside a machine utilizing the reverse proxy, use this URL:

    http://llm-proxy/api/llm/<model-name>/health

    Note: Authorization is not required to access the health check endpoints in any of the scenarios listed above.

    Potential ErrorsPermanent link for this heading

    CUDA is out of memoryPermanent link for this heading

    If the model is deployed on a machine which has not enough GPU memory space, starting a container would cause it to constantly restart. When typing:

    docker logs <container-id or container-name>

    we could see the error “CUDA is out of memory”. In such case, the installation should be done on a machine which has substantial GPU memory space.

    Proxy Service Fails to RunPermanent link for this heading

    If we inspect the status of the llm-proxy service, which is running on the host machine:

    systemctl status llm-proxy

    and the result is something like this:

    llm-proxy.service: Failed to load environment files: Permission denied
    llm-proxy.service: Failed to run 'start-pre' task: Permission denied
    llm-proxy.service: Failed with result 'resources'.

    We could solve this by setting the SELinux to permissive mode:

    setenforce permissive

    To verify that it is permissive type:

    getenforce

    After this we can restart the llm-proxy service and it should be running:

    systemctl restart llm-proxy

    systemctl status llm-proxy

    Download PDF

    • Whitepaper - Mindbreeze InSpire LLM_ On-Premise Deployment Guide

    Content

    • Introduction
    • Prerequisites
    • Setup of Service Account
    • Installation Process
    • Deploying a Large Language Model
    • Appendix

    Download PDF

    • Whitepaper - Mindbreeze InSpire LLM_ On-Premise Deployment Guide