Log in
Textual now supports .pptx files.
For datetime entity types, the synthesis configuration now includes an option to synchronize the shifted values with Structural. In Structural, the columns must use the Day date part. The base seeds and shift ranges in Textual and Structural must also match.
2-series Mastercard numbers are now detected as CREDIT_CARD
Textual now detects “ma’am” and its variants as gender identifiers.
Fixed an issue where Textual could not redact PDFs that contained XFA content.
Fixed blank or incomplete PDF redaction and synthesis output for documents containing inline images.
Custom model entity types can now use multiclass LoRA models, allowing one imported model to identify multiple entity labels while preserving existing single-label model behavior.
Added an opt-in tenant-fair background job scheduling policy for multi-tenant deployments. When enabled, worker capacity rotates across organizations within each job type. This ensures that one organization with a large backlog cannot indefinitely delay work from others. First in first out (FIFO) remains the default, and jobs within each organization retain FIFO order.
Fixed dataset-by-name API requests to return HTTP 404 for missing datasets and improved Python SDK dataset listing reliability during concurrent deletions.
Upgraded the Solar.Py transformers to 5.10.1.
Improved Textual security by updating a machine-learning dependency.
Improved cleanup of expired PDF previews and added optional cleanup on API startup with SOLAR_PDF_PAGE_CACHE_PURGE_ON_STARTUP=true.
You can now filter a dataset file list based on the file scan status (scanning, scanned, or failed).
When users start a bulk re-scan, Textual now displays a confirmation dialog that outlines the impact of running the scan, including the word usage and OCR consumption.
From the file preview, you can now:
- Add manual redactions to Word documents.
- Ignore specific entity instances in text and Word documents.
Fixed style detection for large PDF documents.
Fixed PDF synthesis to preserve visible text colors when documents contain transparent searchable text layers.
Fixed form values disappearing from some de-identified PDF previews and downloaded files.
Fixed an issue where zooming in on a PDF or image file preview clipped the left edge of the page and prevented scrolling to it.
When a user changes their password, Textual now signs the user out all of their active sessions.
Licensed customers can now continue to perform redaction actions after they exceed their allotted word count.
Fixed issues with missing content and white strips in generated output and previews for rotated PDFs that are processed with the V1 PDF synthesis policy.
Download selected dataset files - On the Project files page of a dataset, you can now select individual files and then click Download selected to download only those files into a single ZIP archive. Your selection is preserved as you filter, sort, and page through the file list.
The preview for a PDF file now scrolls continuously through the document's pages. Instead of clicking a page thumbnail, you can scroll from one page to the next. As you scroll, the page indicator, the highlighted thumbnail, and the page shown in the URL follow along. You can still click a page thumbnail or a pagination arrow to navigate the preview to that page.
Improved reliability for concurrent PDF downloads and redactions.
Fixed image redaction so multipage TIFF downloads redact every intended page while preserving the source encoding, color depth, and orientation of every page.
Updated Textual's AWS integrations to AWS SDK for .NET V4 after V3 entered maintenance on March 1, 2026 and reached end-of-support on June 1, 2026. No new Textual environment variable or Helm value is required; before upgrading, deployments must retain only the intended AWS credential source, ensure each client can resolve a Region, use IMDSv2 with metadata response hop limit 2 when Docker relies on instance metadata, and allow regional AWS STS and S3 endpoints.
Added Spanish-, Malay-, and German-aware PDF synthesis that identifies sensitive values in each language and preserves language-specific font styling and text fit.
Added Italian-aware PDF synthesis that identifies sensitive values in Italian and preserves Italian font styling, accents, date conventions, and text fit.
Improved PDF text fitting for accented characters and typographic punctuation by measuring their actual font advances and kerning.
Improved generation performance for text, HTML, and JSON files that have non-standard date formats.
The Tonic MCP server now supports the ability to list and redact dataset files. The tools include options to configure the redaction, such as setting the entity type handling options, allow and block lists of specific values, and settings for specific file types.
Form label mapping - In a form label mapping, you map labels in uploaded PDF forms to entity types. For example, you might map the label Government ID to the SSN entity type. After you set up a mapping, you can attach it to datasets. When Textual detects a mapped form label in a dataset PDF file, it identifies the corresponding value as the mapped entity type.
After a file is rescanned, the entity details popup panels on PDF file previews now correctly display the current original value.
Added support for detecting sensitive values in PDF files that contain Mandarin Chinese (Simplified Chinese).
Added LLM-based sensitivity identification and language-preserving synthesis for typed and scanned PDFs that contain Hong Kong Traditional Chinese, including documents that contain a mix of Chinese and English. Requires that an LLM and Azure Document Intelligence are configured.
Added sensitivity detection and value synthesis for PDF files that contain content in Cantonese.
Fixed blank XFA form values in PDF previews and downloads.
Fixed PDF dataset file previews so that numeric option labels next to form checkboxes are no longer misclassified as phone numbers or displayed with overlapping redaction boxes.
Fixed an issue where, after a custom entity type was enabled or disabled for a dataset, the Scan files to apply changes banner did not display until the page was refreshed. The banner now appears immediately.
Improved the dataset file preview for large text files. Previews now load in small character-based pages with infinite scrolling. This allows large .txt files, including single-line files, to open quickly instead of failing to load.
Added support for Azure when using OpenAI provider.
Fixed an issue that occasionally blocked users from reviewing test files for a model-based entity type.
Using color to identify the entity type handling option - In the dataset file previews for text files only, Textual now highlights detected entities based on the entity type handling option:
Purple for redacted values
Blue for synthesized values
Gray for ignored values
The legend in the file toolbar identifies the colors. From the legend, you can show or hide the highlighting for entity values for each handling type. You can still click the values to change the handling option. Textual remembers these settings across files and sessions.
Confidence score thresholds for custom entity types - For each custom entity type in a dataset, you can now set a minimum detection confidence score. On the dataset details, the new Confidence Thresholds tab displays the distribution of confidence scores for each custom entity type. It includes a slider to select the cutoff. Textual ignores detections that score below the threshold. Those values are not redacted or synthesized. The change is applied immediately and does not require a file rescan. In the file preview and entity catalog, ignored detections remain visible, but are marked as Ignored.
On a dataset's Files page, you can now upload files by dragging and dropping them anywhere on the page. This works both for new datasets and for datasets that already contain files.
Fixed an issue where individual signature detections could not be ignored in PDF previews and downloaded files.
Improved support for synthesizing times.
Fixed an issue where when you added a manual redaction to a dataset file, the Entity Type dropdown only listed built-in entity types. You can now also select any custom entity type that is activated for the dataset.
The Request Auditor now supports Redact API requests for JSON and XML.
Users are now automatically signed out after a configurable period of inactivity.
Textual now provides initial support for an MCP server on both Textual Cloud and self hosted instance. This initial version of the MCP server supports creating datasets and uploading files.
Improved entity detection for Hebrew language
Fixed an issue where an incorrect permission check blocked authorized users from viewing the entity types page.
Updated OpenAPI and transformers dependencies to remediate known vulnerabilities.
Request Auditor - The new Request auditor allows you to review the quality of the entity detection in redaction requests to the Python SDK or REST API. It uses an LLM to detect entities in sample subset of the requests and then compares the LLM results to the original detection results. You then review requests that score at or below a configured threshold to see what was missed in the original detection. Textual automatically removes requests based on a configured expiration. The Request Auditor replaces the Request Explorer.
A new dataset permission, Edit File Redactions, controls whether a user can add or remove manual redactions in a dataset file. The permission is automatically granted to the built-in Editor dataset permission set. Previously, the manual redaction option was controlled by the Edit Dataset Settings permission.
Administrators can now restore access to revoked users.
Fixed an issue where dataset sensitivity information requests timed out while summarizing datasets with large sets of sensitivity results.
Fixed an issue where the Datasets page included links to datasets that the user could not display because they did not have access to them.
Added a sortable Last Updated column to the Entity Types page.
On the Guidelines Refinement page for model-based entity types, the file review now provides a comparison of the established test data values with the values detected by the guidelines version. Values are either true positive (detected in both), false negative (in the established values but not detected by the guidelines version), or false positive (detected by the guidelines but not in the established values).
Fixed issues with rendering synthesized font and currency values in PDFs.
For datasets that connect to Amazon S3, you can now use an assumed role for authentication.
Fixed an issue where an overly restrictive permission check blocke users from activating custom entity types in a dataset.
We've redesigned the entity types page. Custom entity type creation now starts with a wizard-style panel to select the type (regex-based or model-based) instead of a dropdown list of types. For model-based types, the page displays the current status for each step in the creation process, which you can use to jump to a specific step. From the new page, you can also display a read-only list of the built-in entity types.
Fixed an issue with sorting the Entities catalog based on the transformation type.
On the file preview for datasets, you can now display the original and redacted versions of the file side by side.
Dataset audit log - On the dataset details page, the new Audit Log tab tracks new and deleted datasets, changes to the dataset configuration and file list, and downloads of files. Each entry includes the timestamp, user, action type, and action details. Access to the Audit Log tab is controlled by a new dataset permission. From the Audit Log tab, you can download the log.
Improved the performance of PDF processing on textual-ml containers.
Fixed an issue with scrolling on the custom entity types page.
Fixed an issue with tables in .docx files not rendering previews correctly.
Fixed signature detection failures on oversized PDF pages.
Fixed V1 PDF image redaction colorspace failures.
For Okta SSO, Textual can now support the authorization code flow with PKCE. To enable this flow on a self-hosted instance, set SOLAR_SSO_OKTA_USE_PKCE=true. For the Cloud SSO configuration, check the Use PKCE flow checkbox..
Fixed an issue where PDF de-identification could fail while detecting styles for documents with large word payloads.
Improved PDF redaction style detection for documents that reference non-embedded fonts.
Improved PDF table entity detection consistency so sensitive values in tables are classified more reliably across repeated scans.
When you use a Helm chart to install Textual, you can now set a custom database schema for the application database.
Fixed an issue with the scrolling for dataset file previews.
Fixed an issue where the Creator filter for custom entity types did not work.
Improved PDF de-identification font matching for text-backed redactions.
Copying settings from another dataset - On the dataset details page, a new Import Settings option allows you to copy the project and entity type settings from another dataset. You cannot copy settings from a dataset that you do not have access to.
The Date Truncation generator is now available as a synthesis option for all datetime entity types. Previously it was only available for birthdates.
On a self-hosted instance, you can now use environment variables to customize the schema name for the Textual application database. You can also configure whether to automatically create the schema if it does not exist when Textual starts.
Fixed an issue with pagination for model-based custom entity types.
Fixed an issue with the dataset file preview for CSV files.
Fixed an issue with deleting PDF redactions in a guided redaction project.
Added an option to export an individual model-based custom entity type to an encrypted file. When you export the entity type, Textual provides the decryption key for you to copy. You can then import the entity type into another instance. When you import the entity type, you provide the decryption key. The imported entity type is inactive by default, regardless of its status in the original instance.
For zip codes, the synthesis configuration now includes options to truncate zip codes to the first 3 digits, and to replace foreign zip code values with zeroes.
Fixed a guided redaction issue where deleted files were not immediately removed from the file list.
For all entity types, the synthesis configuration now includes an option to provide a single constant replacement value. When you provide a constant value, Textual ignores any entity type-specific synthesis configuration. It does respect mappings of specific original values to replacement values.
Fixed an issue with searching for a user to grant permission set access to.
For the Date of Birth entity type, when you choose to synthesize values, you can now select whether to use the Date Shift Generator, which was previously the only option, or the Date Truncation Generator. For the Date Truncation Generator, Textual always sets the month and day to January 1. If the original year is less than 90 years ago, then Textual keeps the original year. If the original year is 90 or more years ago, then Textual sets the year to the current year minus 89.
For the Age entity type, when you choose to synthesize values, you can now select whether to use the Age shift generator, which was previously the only option, or the Passthrough or group age generator. For Passthrough or group age, if the original age is less than 90, then Textual keeps the original age. If the original age is 90 or older, then Textual sets the age to 90+. This is to meet HIPAA Safe Harbor requirements.
For CSV files in a dataset, from the file preview, you can now configure entity type handling for entire columns. A column might be structured and sensitive, structured but not sensitive, or unstructured. Structured columns contain a single consistent type of value, such as a name, date, or number. For unstructured columns, such as a description or notes, the values can contain multiple entities of different types, and must be scanned individually. Note that this only applies to CSV files that are added to a dataset after this release. Existing CSV files are not affected.
Fixed HTML replacements so that inline breaks such as `
` do not pull adjacent text into phone and email replacement values.
Fixed HTML report mappings so that spanning HTML replacements keep their surrounding markup in report values.
Added an optional phone metadata flag to preserve US prefixes such as `(855)` or `1 (234)` during synthesis.