curiosity

Developers

Extract, enrich
and connect

Raw content is not knowledge. The part number in a scanned report only becomes useful once it is recognized as the same part number the BOM uses.

One part, spelled four ways, joined A part number is extracted from a scanned report. The same part appears as 4471 a, 4471A, 4471-A and P4471A in the report, the maintenance log, the BOM and a supplier notice, and all four join into one part record. SCANNED REPORTP/N 4471 aSCANNED REPORT4471 aMAINTENANCE LOG4471ABOM4471-ASUPPLIER NOTICEP4471APART4471-A
As data lands: Extract, Enrich, Normalize Extract: Text and structure out of files and media. Enrich: Entities recognized, records classified. Normalize: The same identifier, spelled four ways, joined. All inside one as data lands. 01 Extract Text and structure out of files and media 02 Enrich Entities recognized, records classified 03 Normalize The same identifier, spelled four ways,joined

Capabilities

From raw content to connected data

Enrichment runs as records arrive, not as a batch job somebody reruns after noticing the gap.

Entity extraction

Parts, people, systems and obligations recognized in free text and linked to the graph nodes they refer to.

Classification

Records typed on arrival so they land in the right part of the model.

Identifier normalization

The join that makes the rest work: one part, however each system chose to spell it.

Documents and media

Text, tables and structure recovered from files that were never designed to be queried.

Your own pipelines

Custom enrichment steps where the defaults do not know your domain vocabulary.

Incremental

New and changed records are enriched as they arrive.

Questions for developers

The things worth asking first

Can I add my own extractors?

Yes. Add pattern spotters, spotters learned from nodes already in the graph, and trained entity models. In C# you can post-process each entity type, or write a code index that runs on every change to a node type.

Does enrichment use an LLM?

Only when you ask for one. Tokenizing, entity spotting and linking run locally without a model. From code, StructuredAI classifies and scores text with a local model by default, or with an LLM provider you name.

What happens when a source record changes?

When your connector syncs the change, the node is updated in place by its key and queued for every index, so parsing, embeddings and entity links catch up without a full rebuild.

Your data. Your infrastructure.

From raw content to connected knowledge