IRI Voracity is a unified, enterprise-grade data management platform designed for end-to-end data discovery, integration (ETL/ELT), governance, migration, test data management, and AI data preparation. Developed by Innovative Routines International (IRI), Voracity consolidates multiple standalone data tools into a single operational stack powered by the high-performance IRI CoSort manipulation engine.
Core Architecture & Execution Model
Voracity utilizes an in-memory, single-pass processing framework to maximize throughput and eliminate infrastructure overhead:
* In-Memory Single-Pass Engine (CoSort): Instead of staging intermediate data
across temporary database tables or disk storage, Voracity executes multi-
threaded data transformation, sorting, joining, data quality filtering, and masking
simultaneously in-memory.
* IRI Workbench IDE: Jobs are managed and built in IRI Workbench, an Eclipse-based
integrated development environment featuring visual workflow designers, schema
mappers, dialog wizards, and syntax-aware text editors.
* Metadata-Driven 4GL (SortCL): Job definitions use open, human-readable SortCL
(Sort Control Language) text scripts, allowing direct integration into Git
repositories, CI/CD automation pipelines, and modern DevOps tools without
proprietary repository lock-in.
Automated Data Classification & Discovery
Before applying protection or pipelines, Voracity uses a centralized data labeling framework within IRI Workbench:
* Unified Data Classification: Users automatically scan and map corporate data
sources to establish shared data classes and groups tailored to specific privacy
laws (e.g., GDPR, HIPAA) or sensitivity levels.
* Persistent Masking Logic: Once data classification rules are mapped to a specific
data class, those exact rules are reused consistently across all integration,
synthesis, and anonymization workflows, preserving referential integrity.
Key Functional Capabilities
* Test Data Management (TDM) & Privacy (RowGen, FieldShield, DarkShield):
o Synthetic Data Generation (IRI RowGen): Creates referentially intact, highly
realistic synthetic test datasets from scratch using data dictionaries, DDLs, and
custom business logic, removing production privacy risks entirely.
o Advanced Structured Governance (IRI FieldShield): Connects to structured
databases, flat files, ASN.1-compatible call detail records (CDRs), legacy COBOL
files, and Excel spreadsheets. While other engines mask data, FieldShield
uniquely scales to support advanced business logic, conditional rules, and
statistical re-identification (re-ID) risk scoring to calculate and mitigate
information leakage from quasi-identifiers.
o Multi-Source & Semi-Structured Security (IRI DarkShield): Built to scan and
redact PII/PHI hidden in unstructured text, PDFs, logs, and images, as well as
complex semi-structured formats like JSON, XML, NoSQL collections, and Kafka
streams. Additionally, DarkShield is highly capable of executing structured data
masking in relational databases, sharing the same data classification matching
patterns as FieldShield.
* AI, Machine Learning & LLM Data Preparation:
o Feature Engineering & Cleansing: Aggregates, normalizes, standardizes, and
cleanses massive multi-source datasets to feed downstream ML training
pipelines.
o Textual ETL & Compliant RAG Pipelines: Sanitizes enterprise documents,
semi-structured logs, and free-text streams while extracting structured attributes
via textual ETL to build clean, compliant vector databases for Retrieval-
Augmented Generation (RAG) and LLM fine-tuning.
o Synthetic Training Sets: Utilizes RowGen to generate synthetic training data to
balance edge-case distributions and mitigate data scarcity without violating
privacy regulations.
* Real-Time Change Data Capture & Sync (IRI Ripcurrent):
o Tracks transactional log changes across relational engines (Oracle, SQL Server,
MySQL, PostgreSQL) to capture inserts, updates, and deletes in real time.
o Dynamically applies FieldShield data masking rules in-flight during replication
streams to protect sensitive targets.
* Data Integration (ETL/ELT) & Data Quality:
o Connects natively across relational databases, NoSQL engines, cloud object
stores (S3, Azure Blob, GCS), message queues (Kafka, MQTT), and legacy
mainframe files (VSAM, ISAM, EBCDIC).
o Applies validation, data scrubbing, deduplication, and schema standardizations
directly within the ETL flow.
Operational Security & Job Governance (IRI OGS)
To ensure enterprise-level compliance and control across data management workflows, Voracity can be paired with the IRI Operational Governance System (OGS) runtime framework. OGS wraps around SortCL-driven execution jobs to enforce job and data access control policies, verify the integrity of job scripts, and capture granular audit logs, giving compliance and security teams full accountability across all transformation, migration, and masking jobs.
Why Choose Voracity Over Traditional Data Stacks
* Consolidated Tool Footprint: Replaces fragmented licenses for standalone ETL,
data masking, quality, CDC, and test data management software with a unified
runtime and design interface.
* High-Throughput Transformation: Reduces hardware utilization and eliminates
staging bottlenecks via multi-threaded, in-memory processing.
* DevOps & Git-Friendly: Bypasses proprietary binary repositories in favor of
plain-text, version-controllable SortCL scripts.
Who Is the Company Behind IRI Voracity?
-
Seller: IRI
-
Year Founded: 1978
-
HQ Location: Melbourne, US
-
LinkedIn® Page: www.linkedin.com
17 employees on LinkedIn®