Skip to main content

File Search Settings

The settings can be accessed via the Settings button on the search page. Here, you specify which files to process, how much text to extract from them, and how to determine the language of documents and queries. You can also see which part of the files has already been indexed and initiate reindexing.

info

There is no general "Save" button on the page — each change is applied immediately after entering a value, selecting from a list, or toggling a checkbox.

The page and the Settings button are available to administrators with the right to manage the file index, but only those with the right to change general settings can modify values; without this right, the fields are view-only. For more details, see the section “Access to the section”.

Project capabilities​

The Project capabilities block shows what is available to the project. It is for informational purposes only — these values cannot be changed here. Each line is marked with an icon: ✓ — feature available, ✕ — feature not available.

  • Feature included in tariff — whether the feature is included in the project's tariff.
  • Text extraction service available — whether text extraction from files is available. If not, the reason is indicated: service address not specified, service rejected the request, or service is unresponsive.
  • OCR available — whether scan recognition is available.
  • Vector search available — whether vector search is available.

If a feature is not available, the related settings below remain visible but are locked: for example, without the feature in the tariff, file processing cannot be enabled, and without OCR, recognition cannot be configured.

Processing file contents​

The checkbox in the header of the Processing file contents block enables the processing of file contents. By default, processing is turned off. Turning it off stops text extraction from new files, but the already built index is preserved, and searching can still be performed on it.

In this block, you can configure:

  • Formats for extraction — file formats from which text is extracted: txt, md, html, pdf, docx, odt, epub, rtf. All formats are selected by default.
  • Sections whose files are indexed — sections whose files are included in the index: Products, Pages, Blocks, Slides, Templates, Discounts, Forms, Events, User Groups, Users, Admins.
  • Maximum characters per file — how many characters of text are extracted from a single file: from 10000 to 2000000, default is 500000. Text exceeding the limit is not included in the index, and the file is considered truncated. Below the field, an estimate of how many files fit within the allowable index size at the current limit is displayed.
Personal documents

The sections Users, User Groups, and Admins contain documents of clients and administrators. If they are selected, the block warns: the text of these files will become available for search to anyone with search rights and access to the section.

OCR (scan recognition)​

The OCR (scan recognition) block is responsible for recognizing text in scans. The checkbox in the header of the block enables recognition, and the Maximum pages per file field limits the number of pages recognized from a single file: from 1 to 500, default is 50.

If recognition is not available for the project, the checkbox and field are locked, and a warning is displayed in the block. Individual scans can also be recognized manually — in the list of indexed files.

Language of documents and queries​

The Language of documents and queries block determines the language in which documents and search queries are processed.

  • Document language detection — how the language of the document is determined: Automatic (default) or Fixed.
  • Default language — the language of documents: russian, english (default), german, french, or simple.
  • Query language resolution — how the language of the search query is determined: Automatic (default), Fixed, or All corpus languages — searching across all languages present among the indexed files.

Below are two checkboxes, both enabled by default:

  • Allow manual file language override — permission to manually correct the language of an individual file;
  • Search by part of a word — finds occurrences within a word but works noticeably slower.

Index coverage​

The Index coverage block shows the status of the index:

  • how many files are available for search out of the total number;
  • how many files are waiting for the text extraction service;
  • how many files are in an unsupported format;
  • how many files were processed with an error;
  • how many files were truncated due to the character limit;
  • how many files are corrupted;
  • how many files were manually excluded from the search;
  • how many files contain questionable text — this line appears only if such files exist.

Each number is a link to the list of indexed files with the filter already applied. If a document is not found in the search, you can immediately see the reason through these links: format not supported, extraction failed, or text truncated.

Below the counters, the size of the index in megabytes and its share of the allowable volume in percentage is displayed.

warning

If the allowable index volume is exhausted, the block warns: new documents will not be processed, but the already built index continues to be searchable. Reindexing in this case will not help — you need to free up space or increase the volume.

If extraction is paused by the platform operator, the block also informs about this. Such a pause is not lifted by the Processing file contents checkbox; queued tasks are not lost and will continue after the pause is lifted.

Reindexing​

To reprocess files, select from the list What to reindex, which files to process:

  • Only unprocessed — only unprocessed files (default value);
  • Only failed — only those that ended with an error;
  • Truncated and stale — truncated and outdated;
  • Entire corpus — all files.

Then click Rebuild index. A progress bar for reindexing will appear below the button: first, files are queued, then the percentage and number of processed files out of the total are shown. While reindexing is in progress, the button is inactive. If there are no files in the selected slice, the message "Nothing to reindex" is displayed instead of the progress bar. The result of the last reindexing — the percentage and number of processed files out of the total — remains below the button even after it is completed.

warning

If the processing of file contents is turned off in the Processing file contents block, files will be queued but will not be processed until processing is enabled. The Rebuild index button is available only to administrators with the right to manage the file index.