Home
Home
German Version
Support
Impressum
26.4 Release ►

Start Chat with Collection

    Main Navigation

    • Preparation
      • Connectors
      • Create an InSpire VM on Hyper-V
      • Initial Startup for G7 appliances
      • Setup InSpire G7 primary and Standby Appliances
    • Datasources
      • Configuration - Atlassian Confluence Connector
      • Configuration - Atlassian Confluence REST Connector
      • Configuration - Best Bets Connector
      • Configuration - Box Connector
      • Configuration - COYO Connector
      • Configuration - Data Integration Connector
      • Configuration - Database Connector
      • Configuration - Documentum Connector
      • Configuration - Dropbox Connector
      • Configuration - Egnyte Connector
      • Configuration - GitHub Connector
      • Configuration - Google Drive Connector
      • Configuration - GSA Adapter Service
      • Configuration - HL7 Connector
      • Configuration - IBM Lotus Connector
      • Configuration - Jira Connector
      • Configuration - JVM Launcher Service
      • Configuration - LDAP Connector
      • Configuration - Microsoft Azure Principal Resolution Service
      • Configuration - Microsoft Dynamics CRM Connector
      • Configuration - Microsoft Exchange Connector
      • Configuration - Microsoft File Connector (Legacy)
      • Configuration - Microsoft File Connector
      • Configuration - Microsoft Graph Connector
      • Configuration - Microsoft Loop Connector
      • Configuration - Microsoft Project Connector
      • Configuration - Microsoft SharePoint Connector
      • Configuration - Microsoft SharePoint Online Connector
      • Configuration - Microsoft Stream Connector
      • Configuration - Microsoft Teams Connector
      • Configuration - Salesforce Connector
      • Configuration - SCIM Principal Resolution Service
      • Configuration - SemanticWeb Connector
      • Configuration - ServiceNow Connector
      • Configuration - Web Connector
      • Configuration - Yammer Connector
      • Data Integration Guide with SQL Database by Example
      • Indexing user-specific properties (Documentum)
      • Installation & Configuration - Atlassian Confluence Sitemap Generator Add-On
      • Installation & Configuration - Caching Principal Resolution Service
      • Installation & Configuration - Mindbreeze InSpire Insight Apps in Microsoft SharePoint On-Prem
      • Mindbreeze InSpire Insight Apps in Microsoft SharePoint Online
      • Mindbreeze Web Parts for Microsoft SharePoint
      • User Defined Properties (SharePoint 2013 Connector)
      • Whitepaper - Migration of Sites Selected Permissions for the MS SharePoint Online Connector
      • Whitepaper - Migration of Tenant-Wide Permissions for the MS SharePoint Online Connector
      • Whitepaper - Mindbreeze InSpire Insight Apps in Salesforce
      • Whitepaper - Overview of Connectors
      • Whitepaper - Web Connector - Setting Up Advanced Javascript Usecases
    • Configuration
      • CAS_Authentication
      • Configuration - Advanced Configuration for Mail Delivery
      • Configuration - Alerts
      • Configuration - Alternative Search Suggestions and Automatic Search Expansion
      • Configuration - Back-End Credentials
      • Configuration - Chinese Tokenization Plugin (Jieba)
      • Configuration - CJK Tokenizer Plugin
      • Configuration - Collected Results
      • Configuration - CSV Metadata Mapping Item Transformation Service
      • Configuration - Entity Recognition
      • Configuration - Exporting Results
      • Configuration - Filter Plugins
      • Configuration - GSA Late Binding Authentication
      • Configuration - Identity Conversion Service - Replacement Conversion
      • Configuration - InceptionImageFilter
      • Configuration - Index-Servlets
      • Configuration - InSpire AI Chat and Insight Services for Retrieval Augmented Generation
      • Configuration - Item Property Generator
      • Configuration - Japanese Language Tokenizer
      • Configuration - JavaScript Transformer Plugins
      • Configuration - Kerberos Authentication
      • Configuration - Management Center Menu
      • Configuration - Metadata Enrichment
      • Configuration - Metadata Reference Builder Plugin
      • Configuration - Mindbreeze Proxy Environment (Remote Connector)
      • Configuration - Personalized Relevance
      • Configuration - Plugin Installation
      • Configuration - Principal Validation Plugin
      • Configuration - Profile
      • Configuration - Reporting Query Logs
      • Configuration - Reporting Query Performance Tests
      • Configuration - Request Header Session Authentication
      • Configuration - Shared Configuration (Windows)
      • Configuration - Vocabularies for Synonyms and Suggest
      • Configuration of Thumbnail Images
      • Cookie-Authentication
      • Documentation - Mindbreeze InSpire
      • I18n Item Transformation
      • Installation & Configuration - Outlook Add-In
      • Installation - GSA Base Configuration Package
      • JWT Authentication
      • Language detection - LanguageDetector Plugin
      • Mindbreeze Personalization
      • Mindbreeze Property Expression Language
      • Mindbreeze Query Expression Transformation
      • SAML-based Authentication
      • Trusted Peer Authentication for Mindbreeze InSpire
      • Using the InSpire Snapshot for Development in a CI_CD Scenario
      • Whitepaper - AI Chat
      • Whitepaper - Create a Google Compute Cloud Virtual Machine InSpire Appliance
      • Whitepaper - Create a Microsoft Azure Virtual Machine InSpire Appliance
      • Whitepaper - Create AWS 10M InSpire Appliance
      • Whitepaper - Create AWS 1M InSpire Appliance
      • Whitepaper - Create AWS 2M InSpire Appliance
      • Whitepaper - Create Oracle Cloud 10M InSpire Application
      • Whitepaper - Create Oracle Cloud 1M InSpire Application
      • Whitepaper - MMC_ Services
      • Whitepaper - Single Sign-On with Microsoft Entra ID or Active Directory Federation Services
      • Whitepaper - Text Classification Insight Services
    • Operations
      • Adjusting the InSpire Host OpenSSH Settings - Set LoginGraceTime to 0 (Mitigation for CVE-2024-6387)
      • app.telemetry Statistics Regarding Search Queries
      • Blacklisting vulnerable kernel modules esp4, esp6, rxrpc - (Mitigation for CVE-2026-43284 _ DirtyFrag)
      • CIS Level 2 Hardening - Setting SELinux to Enforcing mode
      • Configuration - app.telemetry dashboards for usage analysis
      • Configuration - Usage Analysis
      • Disabling algif_aead_init - (Mitigation for CVE-2026-31431)
      • FAQ - Creating Mindbreeze InSpire Appliances on Hyper Scalers
      • Handbook - Backup & Restore
      • Handbook - Command Line Tools
      • Handbook - Distributed Operation (G7)
      • Handbook - Filemanager
      • Handbook - Indexing and Search Logs
      • Handbook - Updates and Downgrades
      • Index Operating Concepts
      • Inspire Diagnostics and Resource Monitoring
      • Provision of app.telemetry Information on G7 Appliances via SNMPv3
      • Restoring to As-Delivered Condition
      • Whitepaper - Administration of Insight Services for Retrieval Augmented Generation
      • Whitepaper - Insight Workplace
      • Whitepaper - Mindbreeze in Microsoft Teams
      • Whitepaper - Mindbreeze in OpenAI ChatGPT
      • Whitepaper - Mindbreeze InSpire LLM_ A Kubernetes Integration Guide
      • Whitepaper - Mindbreeze InSpire LLM_ On-Premise Deployment Guide
      • Whitepaper - Natural Language Question Answering (NLQA)
      • Whitepaper - Overview AI based Document Parsing, Transcriptions, and Semantic Index Pipeline
      • Whitepaper - Overview of Agentic AI and the Insight Workplace
      • Whitepaper - Using the Mindbreeze InSpire MCP Server
    • User Manual
      • Browser Extension
      • Cheat Sheet
      • iOS App
      • Keyboard Operation
    • SDK
      • api.chat.v1beta.generate Interface Description
      • api.v2.alertstrigger Interface Description
      • api.v2.export Interface Description
      • api.v2.personalization Interface Description
      • api.v2.search Interface Description
      • api.v2.suggest Interface Description
      • api.v3.admin.SnapshotService Interface Description
      • Debugging (Eclipse)
      • Developing an API V2 search request response transformer
      • Developing Item Transformation and Post Filter Plugins with the Mindbreeze SDK
      • Developing Item Transformation Launched Service with Mindbreeze SDK
      • Development of a Query Expression Transformer
      • Development of Insight Apps
      • Embedding the Insight App Designer
      • Export and Integration of Personalization and Analytics Data with External Platforms
      • Java API Interface Description
      • OpenAPI Interface Description
      • SDK Overview
    • Release Notes
      • Release Notes 20.1 Release - Mindbreeze InSpire
      • Release Notes 20.2 Release - Mindbreeze InSpire
      • Release Notes 20.3 Release - Mindbreeze InSpire
      • Release Notes 20.4 Release - Mindbreeze InSpire
      • Release Notes 20.5 Release - Mindbreeze InSpire
      • Release Notes 21.1 Release - Mindbreeze InSpire
      • Release Notes 21.2 Release - Mindbreeze InSpire
      • Release Notes 21.3 Release - Mindbreeze InSpire
      • Release Notes 22.1 Release - Mindbreeze InSpire
      • Release Notes 22.2 Release - Mindbreeze InSpire
      • Release Notes 22.3 Release - Mindbreeze InSpire
      • Release Notes 23.1 Release - Mindbreeze InSpire
      • Release Notes 23.2 Release - Mindbreeze InSpire
      • Release Notes 23.3 Release - Mindbreeze InSpire
      • Release Notes 23.4 Release - Mindbreeze InSpire
      • Release Notes 23.5 Release - Mindbreeze InSpire
      • Release Notes 23.6 Release - Mindbreeze InSpire
      • Release Notes 23.7 Release - Mindbreeze InSpire
      • Release Notes 24.1 Release - Mindbreeze InSpire
      • Release Notes 24.2 Release - Mindbreeze InSpire
      • Release Notes 24.3 Release - Mindbreeze InSpire
      • Release Notes 24.4 Release - Mindbreeze InSpire
      • Release Notes 24.5 Release - Mindbreeze InSpire
      • Release Notes 24.6 Release - Mindbreeze InSpire
      • Release Notes 24.7 Release - Mindbreeze InSpire
      • Release Notes 24.8 Release - Mindbreeze InSpire
      • Release Notes 25.1 Release - Mindbreeze InSpire
      • Release Notes 25.2 Release - Mindbreeze InSpire
      • Release Notes 25.3 Release - Mindbreeze InSpire
      • Release Notes 25.4 Release - Mindbreeze InSpire
      • Release Notes 25.5 Release - Mindbreeze InSpire
      • Release Notes 25.6 Release - Mindbreeze InSpire
      • Release Notes 25.7 Release - Mindbreeze InSpire
      • Release Notes 25.8 Release - Mindbreeze InSpire
      • Release Notes 26.1 Release - Mindbreeze InSpire
      • Release Notes 26.2 Release - Mindbreeze InSpire
      • Release Notes 26.3 Release - Mindbreeze InSpire
      • Release Notes 26.4 Release - Mindbreeze InSpire
    • Security
      • Known Vulnerablities
    • Product Information
      • Product Information - Mindbreeze InSpire - Standby
      • Product Information - Mindbreeze InSpire
    Home

    Path

    Sure, you can handle it. But should you?
    Let our experts manage the tech maintenance while you focus on your business.
    See Consulting Packages

    Whitepaper
    Overview AI based Document Parsing, Transcriptions, and Semantic Index Pipeline

    Executive SummaryPermanent link for this heading

    Mindbreeze InSpire AI-Based Document Parsing and Transcription is an end-to-end ingestion, understanding, and enrichment pipeline that transforms heterogeneous enterprise content – office documents, scanned files, images, audio, videos, etc., - into semantically structured, searchable knowledge. The platform combines Optical Character Recognition (OCR), Vision Language Models (VLMs), Speech-to-Text (STT), layout-aware parsing, Natural Language Processing (NLP), and LLM-driven semantic enrichment to feed a unified AI Search Index and a Retrieval-Augmented Generation (RAG) service.

    The result is a single, consistent representation of enterprise knowledge: every source format is normalized into text with formatting annotations, semantically chunked, enriched with entities and typed facts extracted against configurable schemas, and indexed for high-precision, fully auditable retrieval.

    Architecture OverviewPermanent link for this heading

    Design Philosophy: From Data to KnowledgePermanent link for this heading

    Enterprise content is not uniform in how accessible its information is. Some formats are language-friendly: a .txt file, for example, stores its text and structure explicitly, and extraction is straightforward. Other formats are encoded in ways that resist direct transcription: PDFs — by far the most common enterprise format — store drawing instructions rather than logical text, frequently combined with highly complex layouts (multi-column pages, nested tables, embedded figures, scanned pages with no text layer at all). Audio and video contain no text whatsoever; their information exists only as sound waves and pixels.

    A single extraction technique cannot serve this spectrum. Mindbreeze InSpire therefore takes a layered approach: lightweight native extraction where formats permit it, and an AI-powered extraction layer wherever content is visually, acoustically, or structurally encoded. Everything converges on one canonical representation, after which a Semantic Pipeline performs the deeper work of converting extracted data into indexed knowledge.

    The diagram below represents the dataflow, end to end:


    Content Rendering and AI-Powered ExtractionPermanent link for this heading

    This stage turns arbitrary binary content into a uniform, linked, machine-readable representation. Everything downstream — semantic chunking, fact extraction, indexing, retrieval — consumes its output.

    Input – Binary ContentPermanent link for this heading

    Connectors deliver binary content from the data sources. In practice these spans four broad families:

    • Office and publishing formats – Microsoft Office documents, PDF, and comparable formats.
    • Markup formats – HTML, Markdown, and similar.
    • Multimedia – video, audio, and images.
    • Structured data – JSON, YAML, XML.

    Output – A Sequence of Mindbreeze Semantic Intermediate Representation ItemsPermanent link for this heading

    The result is a sequence of Semantic Intermediate Representation (SIR) items. Each item is a self-contained unit of content carrying:

    • Content – the extracted text or media description, in the canonical representation.
    • Annotations – structural markup, layout roles, captions, semantic tags, and the instruction context under which they were produced.
    • Provenance – the coordinates that locate the item in its source: page number and region for documents, timestamp for audio and video.
    • Links – references to every related artifact: the rendered PDF, the page image, the thumbnail, the originating object in the source system.

    Semantic Intermediate Representation ItemsPermanent link for this heading

    The Mindbreeze SIR Item (mindbreeze.common.Item) is the format for representing indexable content. An Item has special header metadata for identifying the Item such as the uniform document id.

    The Item consists of a sequence of properties (mindbreeze.common.NameValue). A Property consists of a list of values (mindbreeze.common.Value), which can be for instance (but not exclusively):

    • String: A value representing Text
    • Quantity: A numeric value with a unit (such as Timestamp, Byte Size, …)
    • Integer: An integer value
    • Float: A floating-point value
    • List: A list of values
    • Path: A list of values represents a path in a tree (can be used for building hierarchies)
    • Reference: Represents a reference to another Item (For linking items and building graph structures)
    • Tensor: A tensor value (e.g. for representing embeddings)
    • DataBytes: A value for representing blobs such as images
    • AccessControlList: A list of principals and grant information.
    • Item: A nested Item which again consists of a sequence of properties.

    A very powerful part of the Item format is that every Value can have annotations that are again the same type as properties (NamedValue)

    As their name suggest annotations provide a means to annotate values and hence provide additional Information. An annotation always has a name and a type. If only part of the object should be annotated, a marker can be used. For instance, a recognized named entity can be annotated on a text via a textregion entity annotation and the respective marker information. Another example where annotations are used is to preserve the formatting information. That a text has a paragraph or a link can be annotated using html annotations, where the name of the annotation represents the tag (<p>, <a>) and the value of the annotation is an Item that preserves the attributes of the HTML element.

    The Multistep ProcessPermanent link for this heading

    Ingestion is not a single transformation but a composition of Content Processing Services, each with its own responsibilities. Composability is the design goal (e.g., an Office document is rendered to PDF via LibreOffice and then processed by exactly the same path as a PDF that arrived from the data source).

    Linking runs throughout. The links a connector supplies for a source object are enhanced with links computed at each step, so the annotated text, the rendered PDF, the page images, and the thumbnails all remain mutually addressable. This is what allows content to be processed by multimodal LLMs while retaining a path back to the original: an answer derived from a VLM's reading of a chart still resolves to the page it was rendered from.

    Content RenderingPermanent link for this heading

    All incoming documents pass through Content Rendering, which normalizes each source into the representations the AI-powered extraction layer consumes. For example, PDFs are converted into high-fidelity visual renderings of individual pages; video into frames. The conversion is fully configurable.

    The principle is preserve what exists, render what doesn't. Where digital text is present, it is preserved losslessly. Where content is only visual or acoustic, the page images and media streams become the input to AI-powered extraction. Most real documents are a mixture, and Content Rendering produces both representations so the extraction layer can use each where it is authoritative.

    The AI-Powered Extraction LayerPermanent link for this heading

    For formats that resist direct transcription — PDFs, scans, images, audio, and video — Mindbreeze InSpire provides a powerful, adaptable, and evolvable AI-powered extraction layer. It can be operated in two modes:

    • Mindbreeze InSpire LLM – a curated model ecosystem running on-premises, on Kubernetes, or available in Mindbreeze SaaS. The ecosystem can also be deployed entirely on the customer's own infrastructure, keeping content inside the security perimeter.
    • Remote LLMs – any customer defined external model endpoint that supports the OpenAI API or the Hugging Face TGI API.

    Because the layer binds to the OpenAI API rather than to specific models, a single pipeline could handle text, images, audio, and video alike. Concretely, it performs:

    Intelligent Character Recognition (ICR)Permanent link for this heading

    Mindbreeze Inspire supports both OCR and ICR, applying each where it is strongest. Classic OCR recognizes printed characters against a known font set with high speed and accuracy. ICR extends this to handwriting, degraded scans, and irregular or non-standard text – cases where matching against a fixed character set falls short and reading depends more on context and stroke pattern than on fixed shapes.

    Text is recognized directly from the rendered page images rather than from any embedded digital layer, with page coordinates preserved throughout. This means previously unsearchable archives – scanned contracts, handwritten forms, faxed correspondence, photographed documents – become indexable content, with every recognized word still addressable back to its exact position on the page.

    Relevant Links

    Deployment and Usage of the DeepSeek-OCR Model

    Optical Content Processing Permanent link for this heading

    Optical Content Processing is driven by Vision Language Models (VLMs) reading the rendered page image directly – not classic OCR reducing a page to flat text, but a model that understands layout, figures, and visual meaning in context.

    This is what distinguishes it from traditional OCR: a VLM returns text and structure – which characters form a table cell versus a caption, which visual region is a chart and what does it mean. Layout detection and image captioning below are both instances of this same VLM-driven reading, applied to structure and to visual content respectively.

    Partitioning, Object Detection, Layout DetectionPermanent link for this heading

    The same visual models detect sophisticated document structures — multi-column reading order, headings, paragraphs, footnotes, and complex tables with merged or nested cells — and annotate every element. Charts and graphs are detected and routed to visual understanding rather than being lost as opaque images.

    Visual training layout detection can be done using the tool Label Studio which can also be run on the InSpire appliance. The exported data can be used to fine tune the used models such as SmolDocling which is an object and layout detection model used by Docling. One can load documents and tag objects as body text, heading, footer, figure, etc. This results in data that attaches bounding boxes per page (x, y, width, height) with the annotated labels. One can also preload the UI with labels predicted by the model such that one can focus on evaluating and correcting the annotations if needed.

    Vision Language Models (VLMs)Permanent link for this heading

    Vision language models analyze photographs, diagrams, charts, and embedded figures in the context of their surrounding document, generating captions, structured descriptions (e.g., the data series and trends of a chart), and semantic tags that are injected at the image's position in reading order.

    Multimedia Processing (Audio, Video)Permanent link for this heading

    Audio Transcription - Speech to Text (STT)Permanent link for this heading

    Content from audio and video recordings are transcribed into time-aligned, speaker-attributed text; timestamps serve as the acoustic equivalent of page numbers for provenance.

    Relevant Links

    Deploying and Using Speech-To-Text (STT) Models

    Configuration of the OpenAI Proxy

    Video context documentationPermanent link for this heading

    Additional to audio streams visual information is processed via vision models. Using annotations and referencing (Graph structures) the streams are fused along the shared timeline into semantic chapters and visual summaries, so the video's context — what was shown, not only what was said — is documented and indexed.

    Model behavior is governed throughout by an instruction component (prompts, grammars, and schemas), making extraction semantics explicit, reviewable, and reproducible.

    The same API binding that unifies these capabilities also keeps the layer open-ended: state-of-the-art models can be adopted the day they are released — as a new deployment inside InSpire LLM or as a remote endpoint — with no pipeline changes. At the same time, all inference can be gated behind the customer's own hardware, so sensitive documents never leave the security perimeter. Open to the model frontier, closed to data egress.

    Relevant Links

    Deploying and Using Multimodal Models

    Transcription of the Video Stream

    Canonical Output and the Rendering Feedback LoopPermanent link for this heading

    Extraction results are fed back into the Content Rendering stage, where the canonical document representation is updated and the new annotations are stored alongside it. This stage also post-processes output that does not match the expected standard — if the AI layer introduced inconsistencies in the Markdown or HTML, they are detected and corrected here rather than propagating downstream.

    A document carries multiple annotations. The primary ones are:

    • Text with formatting annotations (HTML). The complete text with structural markup preserved.
    • Markdown with page information. A cleaned, portable, rendering that retains page provenance.

    Because rendering consumes its own output, the canonical representation is never a one-shot artifact. As models improve or instructions are refined, documents can be re-processed and their annotations enriched in place — no downstream consumer needs to change.


    Storing SIR Items in the IndexPermanent link for this heading

    Addressing – Uniform Item IDsPermanent link for this heading

    The items are uniquely identified by their category, category instance and the key. The category and category instance are typically used to represent the connector and the connector’s instance. Mindbreeze differentiates between the data source unform item id and the uniform item id of the indexed item. The former is used to keep track of the object in the data source while the latter is used to represent the identity of the object in the index. If a data source object is represented by multiple index items there can be multiple documents having the same data source id.

    Deduplication and Change DetectionPermanent link for this heading

    Before Storing Items in the Index multiple levels of deduplication are performed, such as hashing metadata and contents to identify duplicates. Furthermore, the content hash is used to store same contents only once (content addressed storage).

    If one wants to remove near duplicates, one can use a deduplication plugin that uses locality sensitive hashing (e.g. Minhash). With such a plugin all items that are almost the same (with a configurable similarity threshold) are only indexed once.

    Transactional Index StorePermanent link for this heading

    The Index stores content in a transactional manner, in multiple storage tables and indices. Part of the transactional store is reference symbol to id resolution and cleaning up replaced versions. In the next sections is explained to turn items in the index store into hybrid searchable objects.


    The AI Semantic Index PipelinePermanent link for this heading

    Extraction produces data, but not knowledge. A perfectly transcribed contract is still just text; knowledge emerges when that text is segmented into meaningful units, connected to its structure, enriched with linguistic understanding, and distilled into typed facts. This is the role of the Semantic Pipeline.

    Input – Semantic Intermediate Representation (SIR) Items Permanent link for this heading

    The pipeline consumes the SIR item sequence produced by ingestion, in reading order. It makes no distinction between an item that originated as a digital paragraph, an OCR'd scan region, a VLM chart description, or a speaker turn in a transcript — all arrive in the canonical representation with their annotations, provenance, and links intact.

    This uniformity is what the previous stage exists to guarantee, and it is why the pipeline below can be described once rather than per format.

    Output – Enriched SIR Items with Annotations for Chunks, Facts, and VectorsPermanent link for this heading

    The pipeline is an endomorphic transformation which at every step enriches the items for instance by transforming metadata and adding annotations and references to other items.

    These are then consumed by the next step to incrementally update written to the AI Search Index and served to the RAG Service. Provenance is kept the whole way: every chunk, fact, and vector resolve back through its item to a page and region, or a timestamp – which is what allows an answer to cite the page it came from.

    The Multi-Pass ProcessPermanent link for this heading

    The Semantic Pipeline is not a linear conveyor belt. At any step, data can be fed back to the LLM layer for further passes, and content may take multiple iterations through the steps below before it is indexed.

    A first pass might chunk a document; a second summarize or classify each chunk; a third extract schema facts using those classifications and recognized entities as context; a fourth validate extracted facts against the source text. Each iteration adds a layer of understanding the next can build on — this compositional design is what allows the pipeline to climb from data to knowledge.

    Use Cases Permanent link for this heading

    By the virtue of the generic enrichment and transformation process the Semantic Pipeline can be used for all tasks that can be done without reprocessing the original content data. This also allows reprocessing to being done without reindexing the text.

    • Language detection
    • Lexical analysis – E.g. Tokenization annotations such as non-whitespace word segmentation for Chinese, Japanese, or Korean text
    • Text regions that represent chunked text– the chunk sequence, each chunk retaining the provenance and structural context of the items it spans.
    • Typed data records - schema-conformant facts extracted against user-defined schemas, linked to the chunks and source coordinates they were derived from.
    • Custom Embeddings - vector representations of chunks, and of other units where configured.
    • Information extraction - recognized entities, classifications, key phrases, sentiment, and any LLM-generated summaries, can be annotated and attached to the units they describe.
    • Transformations – Such as Data Masking or Translation of Text/Metadata in other languages

    In the following subsections we will look at use cases in more detail:

    Lexical and Linguistic AnalysisPermanent link for this heading

    There are many built in lexical and linguistic analysis steps that can be configured on the index service configuration settings. The flexibility of the semantic pipeline also allows for custom lexical and linguistic tools can be plugged in easily.

    In addition to adding custom tools also the built-in systems such as machine learning model-based language detection, named entity recognition, compound splitting, sentence segmentation, stop words processing, all can be configured and tuned by adding custom models and catalogs.

    In addition to performing stemming and stop word processing within the semantic pipeline, it is also applied at query time within the query transformation pipeline. The same applies to text transliteration and normalization such as accent and diacritic removal, etc.

    Stop Character ClassesPermanent link for this heading

    Stop character can be used to guide the character level tokenization process.

    TokenizationPermanent link for this heading

    There are multiple tokenization processes happening. For full text search a highly efficient out of the box a highly efficient regular expression tokenization engine (with liner time complexity). Additional semantic tokenization is used for languages such as Chinese, Japanese, Korean (CJK) for instance. If needed custom tokenization plugins can be enabled and added (such as HanNLP, …)

    Language DetectionPermanent link for this heading

    Language detection is done using a very fast neural network for language identification. One can easily plugin in custom language detectors using Item

    Sentence SegmentationPermanent link for this heading

    Sentence segmentation is used to semantically detect sentence boundaries. Which is also the default unit for doing recognizing named entities and performing text segmentation (chunking) for word embedding.

    Custom Analysis PluginsPermanent link for this heading

    If needed custom lexical and linguistic analysis processors can be easily added (such as NLTK, spaCy, Stanza. ..) in the Semantic Pipeline.

    Text Classification and Sentiment Analysis, Named Entity RecognitionPermanent link for this heading

    In addition to prompt based LLM approaches, Mindbreeze InSpire supports efficient language model AI based approaches to NLP tasks such as text classification, sentiment analysis as well as named entity recognition, we build on the Generalist Model for Named Entity Recognition using Bidirectional Transformer (GLiNER and GLiNER2) model architecture. This also enables key phrase extraction.

    There are built-in models available as well as support for training and fine-tuning new models.

    For fine-tuning and training, one needs labelled data is used to adapt as well as train new models. There is also widget support in Insight Apps to create labelled data from interactions such as editable property values for classification, tagging, sentiment, as well as editing entities within the document preview.

    Users can search for entities directly, see them highlighted in previews, and filter results by them. Entities can also feed the embeddings described in the next section, making them findable by meaning as well as by name.

    In addition to machine learned named entity recognition there are also rule based and fuzzy catalog-based solutions. Catalogs can use RDF like formats such as taxonomies, ontologies to fuzzy match and link text to hierarchies and concepts. Pattern based matchers allow syntactical entities to be identified.

    Semantic ChunkingPermanent link for this heading

    With AI document parsing Item annotations represent the semantic units of the documents which can be used in the text segmentation (chunking) step. Furthermore, using token embeddings semantic boundary transition can be detected.  

    Thus, natural topic boundaries and semantic meaning, not fixed character counts. Because chunking operates on the annotated representation, structural boundaries are hard constraints — a chunk never severs a table from its header or a caption from its figure — while embedding-based topic-shift detection places soft boundaries within long sections.

    Every chunk carries its annotations: page, section path, and contained element types. A chunk is therefore addressable in the source document, not merely a span of characters.

    GeocodingPermanent link for this heading

    Within Mindbreeze Insight Apps and Insight Touchpoints geo information is automatically supported if the geo_latitude and geo_longitude properties are present on items. This data is used to display indexed and retrieved items on the map widget, and query items within geographic areas. Another type of mapping is mapping geo coordinates to locations such as country and location names. This allows for instance images that have EXIF GPS latitude and longitude information with filters for guided navigation.

    Geotagging can be implemented by mapping textual address texts via fuzzy catalog enrichment to the representative geo tags. Furthermore, API calls such as Geocoding APIs (Open Source such as https://nominatim.org/, or commercial ones) can be easily integrated using a Script based Semantic Pipeline step.

    Relevant Links

    Development of Insight Apps – map widget

    Metadata Enrichment – Catalog Enrichment

    Instruction and Prompt Based LLM Extraction and EnrichmentPermanent link for this heading

    Using LLM call pipeline processors, predefined and custom prompts can used to perform many tasks such as:

    • Text summarization
    • Text translation
    • Text classification
    • Named Entity recognition
    • Key phrase extraction
    • Sensitive data detection and masking
    • Schema based fact extraction
    • Information extraction in general

    Parts or the complete item is sent to the LLM which then responds with structured output that is then used to transform the item accordingly. Scripting allows for easy, flexible and powerful customizations.

    Smaller focused language models (LMs) can be used to perform tasks using less computation costs.

    Script based Item TransformersPermanent link for this heading

    Mindbreeze InSpire allows you to easily use script-based processor plugins called item transformers. There are also ready-made transformers for LLM interactions, such as calling external LLMs or MCP tools.

    The script can take advantage of pre-made components, including:

    • getMetadataByName: Access existing metadata.
    • sendContentToGenPipeline: Send textual content to an LLM with custom prompts.
    • addMetadatum: Add enriched metadata.

    These components can be used, for example, to perform multiple enrichments for a single document:

    // identify people and roles

    const metadata = sendContentToGenPipeline(input, extractExpertsPrompt);

    // summarize the essence of the content

    Object.assign(metadata, sendContentToGenPipeline(input, summarizePrompt));

    // identify and acronyms and their corresponding full forms

    Object.assign(metadata, sendContentToGenPipeline(input, extractAcronymsPrompt));

    Relevant Links

    Configuration – JavaScript Plugins

    SDK Overview – Extension Points

    Linked Item Contextualization and ValidationPermanent link for this heading

    Mindbreeze InSpire SIR Items are linked items that can use references to build graphs. During Item transformation these newly processed graph nodes can be compared against the existing graph nodes and check if the facts are still relevant.

    Facts are not absolute but also with respect to from the context the facts are linked. A good example is for instance in the health care domain: A patient has risk factors. His clinical documents reference the diagnosis. Yet it is important to match the time of the clinical documents to the risk factors of the patient. E.g. clinical documents of the same patient at a younger age might reference different risk factors than at an older age. In Mindbreeze one can express context filters to facts such that facts also have a temporal dimension.

    Furthermore, using the instruction and prompt based LLM steps in the pipeline one can also use AI to help validate and contextualize facts further.

    Relevant Links

    Mindbreeze Property Expression Language – Paths

    Mindbreeze Property Expression Language – Inverse References

    Mindbreeze Property Expression Language – String References

    Mindbreeze Property Expression Language – Inverse String References

    Mindbreeze Property Expression Language – Transitive Closure

    Metadata Enrichment – Synthesized Metadata

    Language Agnostic SDK Based Item Transformer PluginsPermanent link for this heading

    Mindbreeze uses the Protobuf message via HTTP RPC or GRPC for integrating custom plugins and extensions. This allows for highly efficient custom item transformation plugins. Also many out-of-the-box

    Relevant Links

    SDK Overview – Extension Points

    EmbeddingsPermanent link for this heading

    Chunks are encoded into dense vector representations that place semantically related content near each other in vector space, enabling retrieval by meaning rather than by term overlap. Vectors are stored alongside their chunks in the AI Search Index and are what the RAG Service searches at query time.

    Relevant Links

    Whitepaper - Natural Language Question Answering (NLQA)

    AI-Based Text Classification and Sentiment Analysis with Feedback LoopPermanent link for this heading

    Documents and chunks are classified against customer-defined category schemes, and sentiment is assigned where the content warrants it.

    The feedback loop is what distinguishes this from a one-shot labelling step: corrections are captured and folded back in, so classification behavior improves against the customer's own content rather than remaining fixed at the model's general-purpose defaults. Classifications also feed forward — a chunk's category is available as context to fact extraction, which can select the schema appropriate to the document type.

    Fact and Schema-Based Data ExtractionPermanent link for this heading

    In parallel with chunking and NLP enrichment, the pipeline extracts information against user-defined schemas. Grammars constrain output structure and type — an invoice number must match its pattern; a date must normalize to ISO 8601 — while natural-language prompts define extraction semantics. Grammar-constrained decoding guarantees syntactically valid, machine-consumable values: typed data records rather than best-effort text, stored alongside the chunks.


    Incrementally Updated IndicesPermanent link for this heading

    The previous stages produce knowledge; this one makes it queryable and keeps it current. Indexing is continuous rather than a build step: content enters the index as it is processed, and re-processed content replaces its earlier version in place.

    Input – Enriched Semantic Intermediate Representation ItemsPermanent link for this heading

    The stage consumes the output of the semantic pipeline: chunks with their annotations, typed fact records, embeddings, and enrichment annotations, each still carrying the provenance and links established during ingestion.

    Items arrive continuously as documents are processed, and the same item may arrive again — when a connector reports a change at the source, or when the rendering feedback loop re-processes a document under improved models or refined instructions. The stage treats both cases identically.

    Output – Incrementally Updated Searchable IndicesPermanent link for this heading

    The stage maintains three complementary representations of the same content:

    • Full-text index – terms, tokens, and their positions, supporting keyword and phrase retrieval, filtering over aggregatable metadata, and the query language.
    • Vector index – chunk embeddings, supporting similarity search and semantic retrieval by meaning rather than term overlap.
    • Graph index – the reference structure between items: entities, links, and the relationships resolved during processing, supporting traversal from a document to what it references and back.

    Incremental UpdatePermanent link for this heading

    Documents are inserted, replaced, and deleted individually. Replacement is detected by comparing modification date and content or metadata checksums, so unchanged content is not re-indexed, and a document that changes is updated in place where the change permits it.

    This is what makes the earlier stages' feedback loop practical. A model improvement that enriches annotations across a corpus does not require a rebuild — re-processed documents flow back through the same insertion path and replace their predecessors, and the index remains queryable throughout.

    Graph Index RepresentationPermanent link for this heading

    In Mindbreeze InSpire, items are represented as linked data structures. This allows for the creation of complex relationships between indexed objects, enabling powerful graph-based retrieval. There are two primary methods for modeling these links:

    1. Reference Value Lookup

    Reference value lookups make relationships explicit by using specific reference values to connect items. This method is typically used to represent structural data source references, including:

    • Hierarchies: Modeling the relationship between files and folders.
    • Content Linking: Establishing links between contents (e.g., when indexing ontologies).
    • Security: Implementing Access Control List (ACL) inheritance.
    • Object Fragmentation: Linking divided objects. During parsing and extraction, a single data source object may be represented as multiple index items (e.g., the rendered PDF, structured text and thumbnails).

    2. String Value Lookup

    String value lookups allow any text string to serve as a reference to another item. This is particularly useful for linking objects, even if there was no explicit reference value set during indexing. Typical use-cases include:

    • Product IDs
    • Invoice numbers
    • E-Mail addresses
    • …

    Capabilities of Graph Retrieval

    Mindbreeze InSpire has the ability to seamlessly combine graph navigation with standard retrieval methods. Key capabilities include:

    • Integrated Authorization: Graph traversals are efficiently filtered via precomputed ACL information, ensuring users only see linked data they are authorized to access.
    • Hybrid Querying with Sub-Queries: Support for recursive sub-queries that combine full-text and semantic search with graph navigation. For example, this can be used to identify subject matter experts by identifying authors who have frequently published content on a specific topic.
    • Computed Properties: Via the Mindbreeze Property Expression Language, properties of linked objects can be retrieved. This approach utilizes the efficient graph representation to consolidate data at query time, providing use-case-specific results without the overhead of reprocessing source documents.
    • Transitive Closures: Ability to perform multi-step traversals to find distant relationships.
    • Bidirectional Traversal: While links are directed, the system can invert references (Reverse Traversal) to find all items pointing to a specific target. For example, if a document references an author, a reference inversion can be used to find all documents written by that specific author.

    Relevant Links

    Mindbreeze Property Expression Language – Paths

    Mindbreeze Property Expression Language – Inverse References

    Mindbreeze Property Expression Language – String References

    Mindbreeze Property Expression Language – Inverse String References

    Mindbreeze Property Expression Language – Filter Annotations

    Mindbreeze Property Expression Language – Transitive Closure

    api.v2.search Interface Description – Computed Properties

    Documentation – Mindbreeze InSpire – Sub Query Expression

    Hybrid Search

    Full-text Index RepresentationPermanent link for this heading

    For performing full-text search, the tokenized text stream is inverted using terms and N-grams. The inverted index divides documents in zone and is aware of positions. To optimize the construction of vector indices, log structured merge is used.

    The retrieval scoring uses the InSpire relevance model and is able to make use of optional terms (a form of weak and) where the missing terms of each individual results are returned in the search response. Additional to BM25 for relevance the position information of the inverted index is used to compute the proximity such that terms that appear close to each other get more weight. Many more parameters are usable.

    As part of the semantic pipeline compound words are split by using split probabilities for the text language. For other sub word pattern retrieval, the N-gram inverted index is used. If regex matching is desired it can be activated on a zone level.

    Vector Index RepresentationPermanent link for this heading

    The Vector Index use a bucketed cluster architecture so that queries are performed on a cluster of documents to scale and shortcut whole clusters. To optimize the construction of vector indices, log structured merge is used. The Vector Index has an IVF representation. During vector construction additionally to the embeddings vector id, the document ids containing the vectors are natively embedded in the vector index format. The similarity search component is optimizing the vector comparisons also that only vectors that are retrievable have to be compared. This is especially effective when using search constraints or a combination of full text search with vector search, or when only part of the document collection is accessible due to access restrictions.

    Vector search can be done in multiple regions. These regions can also be combined. During index construction Mindbreeze InSpire efficiently builds embeddings for multiple regions such as default (text segmentation on sentence level with overlaps), large (combining multiple default ranges), page (if page annotations are available the granularity are whole pages), document (the complete document vector). By default, the larger regions use pooling strategies to combine the individual element vectors.

    The answer text will also match the search region. Default returns the region of the chunks with overlaps. Large will return all the regions that were accumulated for the large region (configurable), page will return the complete text of the page, and document the whole documents text. Answers optionally contain structured information such as tables and paragraphs, if available as annotations on the text and other annotations such as entities that have been highlighted.

    Larger regions therefore are used in retrieval augmented generation (RAG) pipelines to pass more context to the LLM for generation. The citation feature in RAG will then provide only the effectively cited text as sources to the response.


    Hybrid SearchPermanent link for this heading

    Unlike traditional hybrid search implementations that rely on Rank Fusion algorithms to merge separate ranked lists, Mindbreeze InSpire utilizes a tightly integrated retrieval engine. In our architecture, Graph, Vector, and Full-text retrieval are unified into a single document hit. This allows a singular, uniform relevance model to compute the final rank score in one pass.

    This deep integration eliminates the „black box“ nature of rank fusion and provides significantly more granular control over result precision. The final relevance score is based on the following:

    • Semantic Intelligence: Vector similarity and similarity weights for reranking.
    • Lexical Factors: Term frequency, document frequency, term proximity, and term inverse zone frequency.
    • Freshness: Incorporating recency to ensure timely information.
    • Weighted Tuning: Granular boosting applied at the zone, term, and document levels, as well as specific (reranked) answer document boostings.

    Integrating these factors into a unified scoring algorithm allows Mindbreeze InSpire to optimize for both semantic recall and lexical precision, bypassing the data degradation often associated with the fusion of independent result sets.

    Relevant Links

    Mindbreeze Query Expression Transformation – Relevance Factors

    Documentation – Mindbreeze InSpire – Sentence Transformation

    Lookup Support – Entity ResolutionPermanent link for this heading

    The index has structures to automatically build string-based lookup targets (also reverse) given reference values (links in the intermediate representation directly) as well as string values such as identified entities.

    A reference value is a link between items that is known by the uniform id of the indexed item (e.g. the key or data source key).

    In the case of string references any property value (also identified in the semantic pipeline e.g. via LLM calls) can be used as a string lookup reference target. In this case this could be e.g. the name of an entity or an email address or an invoice id. These lookup values are very efficient and therefore one can use it via query to find all referenced objects by the power of the search using the effective permissions of the user doing the query.

    Hence based on the graph, entities are linked as references via the Semantic pipeline yet the final resolution can be done highly efficient using hybrid graph retrieval dynamically given the current graph information and the context of the user. One use case is that the entity targets might have a dynamic lifecycle – entities are later added, updated or removed, another use case of dynamic resolution is that one is only interested in entities being referred to from content that is of relevance for the user.

    Relevant Links

    Graph Index Representation

    Linked Item Contextualization and Validation

    Mindbreeze Property Expression Language – String References

    Mindbreeze Property Expression Language – Inverse String References

    Mindbreeze Property Expression Language – Filter Annotations

    ConsumersPermanent link for this heading

    The indices serve retrieval directly, and ground the RAG Service's generative answers in retrieved chunks. Because every chunk carries page numbers, section paths, or timestamps, every generated statement is traceable to its exact location. Knowledge, once built, is consumed with full auditability.

    Further InformationPermanent link for this heading

    Topic

    Link

    Ingestion and Connectors

    Connector overview (500+ data sources)

    https://help.mindbreeze.com/en/doc/Connectors/index.htm

    Filter Services and content processing

    https://help.mindbreeze.com/en/index.php?topic=doc/Configuration---Filter-Plugins/index.htm

    Thumbnail and image page rendering

    https://help.mindbreeze.com/en/doc/Configuration-of-Thumbnail-Images/index.htm

    Image classification filter

    https://help.mindbreeze.com/en/doc/Configuration---InceptionImageFilter/index.htm

    Structured data ingestion

    https://help.mindbreeze.com/en/doc/Configuration---Data-Integration-Connector/index.htm

    Custom item transformation (SDK)

    https://help.mindbreeze.com/en/doc/Developing-Item-Transformation-and-Post-Filter-Plugins-with-the-Mindbreeze-SDK/index.htm

    Index Service Settings

    https://help.mindbreeze.com/en/doc/Documentation---Mindbreeze-InSpire/index.htm#index-service-settings


    Semantic Pipeline

    Language detection

    https://help.mindbreeze.com/en/index.php?topic=doc/Language-detection---LanguageDetector-Plugin/index.htm

    Named Entity Recognition (NER)

    https://help.mindbreeze.com/en/index.php?topic=doc/Documentation---Mindbreeze-InSpire/index.htm#named-entity-recognition-ner

    Compound Splitting

    https://help.mindbreeze.com/en/index.php?topic=doc/Documentation---Mindbreeze-InSpire/index.htm#compound-splitting

    Rule-based entity recognition

    https://help.mindbreeze.com/en/index.php?topic=doc/Configuration---Entity-Recognition/index.htm

    Metadata enrichment and custom annotations

    https://help.mindbreeze.com/en/index.php?topic=doc/Configuration---Metadata-Enrichment/index.htm

    Embeddings, sentence transformation, NLQA

    https://help.mindbreeze.com/en/index.php?topic=doc/Whitepaper---Natural-Language-Question-Answering-NLQA/index.htm

    Text classification

    https://help.mindbreeze.com/en/index.php?topic=doc/Whitepaper---Text-Classification-Insight-Services/index.htm

    Property expression language

    https://help.mindbreeze.com/en/index.php?topic=doc/Mindbreeze-Property-Expression-Language/index.htm

    Scripting Engine

    https://help.mindbreeze.com/en/index.php?topic=doc/Configuration---JavaScript-Transformer-Plugins/index.htm#introduction

    Index and Retrieval

    Index operating concepts, incremental update, reinversion

    https://help.mindbreeze.com/en/index.php?topic=doc/Index-Operating-Concepts/index.htm

    Query language and query expression transformation

    https://help.mindbreeze.com/en/doc/Mindbreeze-Query-Expression-Transformation/index.htm

    Relevance and personalization

    https://help.mindbreeze.com/en/doc/Configuration---Personalized-Relevance/index.htm

    RAG and Generative AI

    Insight Services for RAG

    https://help.mindbreeze.com/en/doc/Configuration---InSpire-AI-Chat-and-Insight-Services-for-Retrieval-Augmented-Generation/index.htm

    RAG Administration

    https://help.mindbreeze.com/en/index.php?topic=doc/Whitepaper---Administration-of-Insight-Services-for-Retrieval-Augmented-Generation/index.htm

    AI Chat

    https://help.mindbreeze.com/en/index.php?topic=doc/Whitepaper---AI-Chat/index.htm

    Chat generation API

    https://help.mindbreeze.com/en/index.php?topic=doc/apichatv1betagenerate-Interface-Description/index.htm

    InSpire LLM

    Mindbreeze InSpire LLM: A Kubernetes Integration Guide

    https://help.mindbreeze.com/en/index.php?topic=doc/Whitepaper---Mindbreeze-InSpire-LLM-A-Kubernetes-Integration-Guide/index.htm

    Mindbreeze InSpire LLM: On-Premise Deployment Guide

    https://help.mindbreeze.com/en/index.php?topic=doc/Whitepaper---Mindbreeze-InSpire-LLM-On-Premise-Deployment-Guide/index.htm

    Download PDF

    • Whitepaper - Overview AI based Document Parsing, Transcriptions, and Semantic Index Pipeline

    Content

    • Executive Summary
    • Architecture Overview
    • Content Rendering and AI-Powered Extraction
    • Storing SIR Items in the Index
    • The AI Semantic Index Pipeline
    • Incrementally Updated Indices

    Download PDF

    • Whitepaper - Overview AI based Document Parsing, Transcriptions, and Semantic Index Pipeline