Analysts now describe the semantic layer as critical enterprise infrastructure by 2030: a universal, machine-readable description of what an organisation knows, sitting between raw data and every application that consumes it. The forecast is right. What it understates is how few enterprises have anything resembling one today, and how much of their knowledge is locked in documents that no semantic layer currently reaches.

What a semantic layer actually is

A semantic layer is not a data catalogue and not a vector index. It is an explicit model of the entities that matter to a business, the relations between them and the rules that govern them, populated from real data and kept traceable to its sources. In a bank, the entities are borrowers, guarantors, properties, contracts, covenants, payments and regulations. The relations are who owns what, who guarantees whom, which clause applies to which exposure. The rules are the ontology: what a valid mortgage is, what an event of default means, how a cadastral identifier is structured.

Once that model exists, an agent no longer answers from text similarity. It answers from structure: this value is the residual capital of this loan, secured by this property, appraised by this expert on this date, with this page as evidence.

From documents to entities

The hard part is population. Most of an institution's knowledge is unstructured: 130 million pages of loan files at one servicer, 10 million company documents at a national registry, tens of thousands of investor reports at an asset manager. Building the semantic layer means reading all of it with domain-tuned models, classifying every document, extracting entities and objects rather than isolated fields, and linking them into the graph with a confidence score and a source.

We organise this population as medallion layers. A bronze layer holds the physical documents and their layout structure. A silver layer holds validated entities and relations. A gold layer holds the business objects processes actually use: a complete borrower profile, a reconciled portfolio position, a regulatory obligation mapped to controls.

Cadastral triples extracted from a court appraisal and reconciled against the property register.
Cadastral triples extracted from a court appraisal and reconciled against the property register.

Why regulated industries go first

Banks, insurers and public registries are adopting semantic layers ahead of everyone else for a simple reason: they cannot use an answer they cannot explain. A GraphRAG query that traverses the ontology and returns the path it followed is admissible in a credit committee and defensible in front of a supervisor. A probabilistic paragraph is not. The same property that makes semantic layers expensive to build, their explicitness, is what makes them mandatory in regulated environments.

Ten years of research, in production

Altilia's semantic technology comes from a decade of work at the Italian National Research Council and the University of Calabria: published research on knowledge graphs, document layout analysis and information extraction, a US patent on semantic document understanding, and a visual question answering programme funded to push the limits of reading complex layouts. That research is what runs today at Tier-1 institutions, not as a prototype but as the layer their agents reason over.

“A semantic layer is the difference between an agent that sounds right and an agent that can show you why it is right.”
Ermelinda Oro, Co-founder

How to begin without a five-year programme

Do not start with a company-wide ontology. Start with one process whose documents you already understand, model the twenty entities that matter to it, and let agents populate the graph while they deliver value. The onboarding ontology becomes the lending ontology; the lending ontology becomes the collateral ontology. Each process you automate leaves behind a piece of the semantic layer the next one needs.

Key takeaways

  • A semantic layer models entities, relations and rules, populated from real data and traceable to sources.
  • Most enterprise knowledge sits in unstructured documents that a semantic layer must read to be complete.
  • Regulated industries adopt first because explicit structure is the only defensible basis for automated decisions.
  • Build it process by process; every automated workflow extends the graph.

Curious how this would work on your documents?