Transforming Text into Structured Knowledge

Books preserve knowledge through narrative.

Dark Mage’s Archive preserves knowledge through structure.

The Semantic Extraction process transforms written material into interconnected semantic entities while preserving the meaning and context of the original source.

Rather than treating a document as a single block of text, the archive identifies the individual concepts that together form its knowledge.


Beyond Text Recognition

Semantic extraction is fundamentally different from text recognition.

Recognizing words is only the first step.

Understanding what those words represent is the real objective.

The extraction process identifies meaningful concepts such as procedures, items, goals, principles, rules, properties, and many other semantic entities, together with the relationships that connect them.

The result is structured knowledge rather than searchable text.


Preserving Meaning

Every source contains information at multiple levels.

A paragraph may describe a procedure.

Within that procedure are individual actions.

Those actions involve specific items.

The items possess properties.

The procedure pursues one or more goals.

Underlying the entire process may be principles or rules that explain why the procedure exists.

Semantic extraction separates these concepts while preserving the relationships that give them meaning.


A Structured Process

The extraction process follows a consistent methodology designed to organize knowledge while respecting the structure of the original source.

Rather than copying documents into a database, information is analyzed, classified, and represented as interconnected semantic entities.

This structured approach allows knowledge from many independent sources to coexist within a common framework.


Context Matters

Meaning depends on context.

The same word may represent different concepts in different traditions or historical periods.

For this reason, semantic extraction considers the surrounding context rather than isolated words alone.

The objective is to preserve meaning rather than merely identify vocabulary.


Relationships Are Extracted Together

Knowledge does not consist only of entities.

Relationships are equally important.

Whenever meaningful connections exist between concepts, they are preserved as part of the extraction process.

These relationships later become the paths through which visitors explore the knowledge graph.

Without relationships, individual entities would lose much of their value.


Designed for Consistency

A consistent extraction methodology allows knowledge originating from different books and traditions to be represented within the same semantic framework.

Although individual sources may differ greatly in style, terminology, or historical background, the resulting knowledge remains organized according to common semantic principles while preserving each source’s original context.


Human-Guided Validation

Technology assists the extraction process, but consistency requires careful validation.

Extracted knowledge is reviewed according to clearly defined semantic rules that help preserve accuracy, minimize ambiguity, and reduce unsupported interpretations.

The objective is not to reinterpret historical material, but to represent it faithfully within the knowledge model.


Building a Living Archive

Every successful extraction enriches the archive.

New entities expand the knowledge graph.

New relationships create additional paths of discovery.

New sources contribute fresh perspectives while strengthening connections with existing knowledge.

Over time, individual books become part of a much larger interconnected network that can be explored as a whole rather than read one document at a time.