ISO/IEC TR 5812:2026
(Main)Information technology — Semantic metadata support in office documents
General Information
- Abstract
This document outlines: descriptions of use cases for identified application domains that benefit from additional semantic metadata support in office documents; descriptions of semantic metadata requirements from experts of identified application domains; specifications of possible technical strategies for associating domain-specific semantic metadata with office documents entities. This document applies to users of office documents, for example, publishers, librarians, archivists, data scientists and other document users, as well as office software developers and service system integrators. This document focuses on the methods of supporting various kinds of metadata, rather than on metadata itself.
- Status
- Published
- Publication Date
- 01-Oct-2026
- Technical Committee
- ISO/IEC JTC 1/SC 34 - Document description and processing languages
- Drafting Committee
- ISO/IEC JTC 1/SC 34 - Document description and processing languages
- Current Stage
- 6060 - International Standard published
- Start Date
- 02-Oct-2026
- Completion Date
- 02-Oct-2026
Overview
ISO/IEC TR 5812:2026 addresses semantic metadata support in office documents and explains how office document ecosystems can better support machine-readable meaning in editable content. The document is aimed at users of office documents, including publishers, librarians, archivists, data scientists, office software developers, and service system integrators.
The technical report focuses on the methods of supporting metadata, rather than defining the metadata itself. It outlines use cases across multiple domains, describes metadata requirements from subject matter experts, and presents possible technical strategies for associating domain-specific semantic metadata with office document entities. This makes the standard highly relevant for organizations working with OOXML and ODF document formats.
Key Topics
Semantic metadata in office documents
- Supports interoperability, search, and machine interpretation
- Helps represent content and context in a structured way
Use cases across communities
- Publishing workflows
- Libraries and archives
- Machine learning and natural language processing
- Collaborative editing environments
Technical strategies
- Associating metadata with document content
- Supporting metadata inside or outside the document
- Preserving metadata during editing and transformation
Representation and interoperability
- References common metadata approaches such as RDF
- Considers vocabularies and ontology-based descriptions
- Supports reuse of metadata across different systems and workflows
Lifecycle value
- Improves information discovery and retrieval
- Supports document organization and presentation
- Assists content analysis and long-term preservation
Applications
ISO/IEC TR 5812:2026 is practical for organizations that need to manage office documents as structured knowledge assets. Common application areas include:
Publishing
- Tagging titles, authors, subjects, sections, and other descriptive elements
- Supporting discovery, indexing, and downstream typesetting workflows
Libraries and archives
- Enhancing cataloguing, classification, and preservation
- Helping maintain meaningful document context over time
Data science and AI
- Annotating documents for retrieval, classification, and analysis
- Supporting traceability for generated or processed content
Office software development
- Improving metadata capture, preservation, and interchange
- Enabling semantic support without replacing existing document workflows
Document conversion workflows
- Preserving metadata during transformations to PDF, web formats, or other outputs
Related Standards
This technical report is closely connected with well-established document and metadata standards, including:
- ISO/IEC 29500 - Office Open XML document formats
- ISO/IEC 26300 - OpenDocument Format (ODF)
- ISO/IEC 11179-1 - Metadata registries
- ISO 19115 - Metadata for geographic information
- ISO/IEC 22989 - Artificial intelligence concepts and terminology
- ISO/IEC 23053 - Framework for AI systems using machine learning
- RDF and RDFa - Semantic web and embedded metadata technologies
For organizations seeking better office document metadata management, semantic annotation, and interoperable document processing, ISO/IEC TR 5812:2026 provides valuable guidance for planning future-ready document workflows.
Get Certified
Connect with accredited certification bodies for this standard

BSI Group
BSI (British Standards Institution) is the business standards company that helps organizations make excellence a habit.

NYCE
Mexican standards and certification body.
Sponsored listings
Frequently Asked Questions
ISO/IEC TR 5812:2026 is a technical report published by the International Organization for Standardization (ISO). Its full title is "Information technology — Semantic metadata support in office documents". This standard covers: This document outlines: descriptions of use cases for identified application domains that benefit from additional semantic metadata support in office documents; descriptions of semantic metadata requirements from experts of identified application domains; specifications of possible technical strategies for associating domain-specific semantic metadata with office documents entities. This document applies to users of office documents, for example, publishers, librarians, archivists, data scientists and other document users, as well as office software developers and service system integrators. This document focuses on the methods of supporting various kinds of metadata, rather than on metadata itself.
This document outlines: descriptions of use cases for identified application domains that benefit from additional semantic metadata support in office documents; descriptions of semantic metadata requirements from experts of identified application domains; specifications of possible technical strategies for associating domain-specific semantic metadata with office documents entities. This document applies to users of office documents, for example, publishers, librarians, archivists, data scientists and other document users, as well as office software developers and service system integrators. This document focuses on the methods of supporting various kinds of metadata, rather than on metadata itself.
ISO/IEC TR 5812:2026 is classified under the following ICS (International Classification for Standards) categories: 35.240.30 - IT applications in information, documentation and publishing. The ICS classification helps identify the subject area and facilitates finding related standards.
ISO/IEC TR 5812:2026 is available in PDF format for immediate download after purchase. The document can be added to your cart and obtained through the secure checkout process. Digital delivery ensures instant access to the complete standard document.
Standards Content (Sample)
Technical
Report
ISO/IEC TR 5812
First edition
Information technology — Semantic
2026-10
metadata support in office
documents
Technologies de l'information — Support de métadonnées
sémantiques dans les documents bureautiques
Reference number
© ISO/IEC 2026
All rights reserved. Unless otherwise specified, or required in the context of its implementation, no part of this publication may
be reproduced or utilized otherwise in any form or by any means, electronic or mechanical, including photocopying, or posting on
the internet or an intranet, without prior written permission. Permission can be requested from either ISO at the address below
or ISO’s member body in the country of the requester.
ISO copyright office
CP 401 • Ch. de Blandonnet 8
CH-1214 Vernier, Geneva
Phone: +41 22 749 01 11
Email: copyright@iso.org
Website: www.iso.org
Published in Switzerland
© ISO/IEC 2026 – All rights reserved
ii
Contents Page
Foreword .iv
Introduction .v
1 Scope . 1
2 Normative references . 1
3 Terms and definitions . 1
4 Document semantic metadata . 1
5 Office document communities . 2
6 Typical use cases . 2
6.1 Themes .2
6.1.1 Information discovery and retrieval .2
6.1.2 Document organization and presentation.3
6.1.3 Document content analysis and processing .3
6.1.4 Long-term document preservation .4
6.2 General questions to consider.5
7 Representation of semantic metadata . 5
8 Metadata used in office documents . 6
9 Storing metadata in OOXML/ODF . 8
10 Associating semantic metadata with content in OOXML/ODF .10
11 Application examples . .11
12 Future work on document format standards . 14
Bibliography .15
© ISO/IEC 2026 – All rights reserved
iii
Foreword
ISO (the International Organization for Standardization) and IEC (the International Electrotechnical
Commission) form the specialized system for worldwide standardization. National bodies that are
members of ISO or IEC participate in the development of International Standards through technical
committees established by the respective organization to deal with particular fields of technical activity.
ISO and IEC technical committees collaborate in fields of mutual interest. Other international organizations,
governmental and non-governmental, in liaison with ISO and IEC, also take part in the work.
The procedures used to develop this document and those intended for its further maintenance are described
in the ISO/IEC Directives, Part 1. In particular, the different approval criteria needed for the different types
of ISO documents should be noted. This document was drafted in accordance with the editorial rules of the
ISO/IEC Directives, Part 2 (see www.iso.org/directives or www.iec.ch/members_experts/refdocs).
ISO and IEC draw attention to the possibility that the implementation of this document may involve the
use of (a) patent(s). ISO and IEC take no position concerning the evidence, validity or applicability of any
claimed patent rights in respect thereof. As of the date of publication of this document, ISO and IEC had not
received notice of (a) patent(s) which may be required to implement this document. However, implementers
are cautioned that this may not represent the latest information, which may be obtained from the patent
database available at www.iso.org/patents and https://patents.iec.ch. ISO and IEC shall not be held
responsible for identifying any or all such patent rights.
Any trade name used in this document is information given for the convenience of users and does not
constitute an endorsement.
For an explanation of the voluntary nature of standards, the meaning of ISO specific terms and expressions
related to conformity assessment, as well as information about ISO's adherence to the World Trade
Organization (WTO) principles in the Technical Barriers to Trade (TBT), see www.iso.org/iso/foreword.html.
In the IEC, see www.iec.ch/understanding-standards.
This document was prepared by Joint Technical Committee ISO/IEC JTC 1 Information technology,
Subcommittee SC 34, Document description and processing languages.
Any feedback or questions on this document should be directed to the user’s national standards
body. A complete listing of these bodies can be found at www.iso.org/members.html and
www.iec.ch/national-committees.
© ISO/IEC 2026 – All rights reserved
iv
Introduction
Office document refers to the electronic document created, edited or processed by office applications (e.g.
word processors, spreadsheets, presentation software), which adheres to structured markup standards
for interoperability [see ISO/IEC 29500 (series)]. OOXML [see ISO/IEC 29500 (series)] and ODF (see
ISO/IEC 26300) are widely used office document formats.
NOTE 1 There are other document types, such as PDF, EPUB and HTML, that incorporate semantic information to
which the methods proposed in this document apply; however, they fall outside the scope of this document.
Documents are becoming more intelligent, which requires that their content be understood by both humans
and machines. Adding tags to markup the semantic content of a document, or incorporating semantic
information into a document through tags, is an effective way to achieve human- and machine-readable
documents. In this document, the document semantic information expressed through tags is referred to
as semantic metadata (abbreviated as metadata). The goal of semantic metadata is to provide a clear and
accurate description of the content and context of a document, in an interoperable and accurate way. These
semantic metadata can be defined in terms of industry vocabularies, resulting in normative metadata that
can be shared across multiple application scenarios.
Traditionally, metadata were primarily used in web pages and fixed-format documents. This usage was
due to the static nature of these documents, which were rarely modified and were suitable for long-term
preservation, making metadata essential for locating key content within these documents. In contrast,
support for metadata in editable documents is limited. It has been generally believed that since editable
documents serve as intermediaries for editing, incorporating metadata is unnecessary.
Furthermore, there is a belief that metadata become unstable as documents undergo editing. However,
recent findings indicate that there are billions of editable documents worldwide today, which are more
numerous and contain richer semantic information compared to fixed-layout documents. These editable
documents are valuable knowledge resources, making it sensible to add metadata to them. Additionally, the
metadata annotated by the author during the document editing process can accurately reflect the semantics
of the document.
To this end, this document proposes semantic support methods for office documents.
NOTE 2 Common office document components include word processing documents, spreadsheets and
presentations. The methods presented in this document generally apply to spreadsheets and presentations as well.
Due to space limitations, the examples provided in this document focus primarily on word processing documents.
NOTE 3 The internal objects in the composite documents based on OOXML and ODF sometimes contain their own
metadata and support for such metadata is also beyond the scope of this document.
There are hundreds of industry vocabularies that exist today, each a potential source of metadata. Typically,
it is up to different user communities to decide which vocabulary to use metadata from.
© ISO/IEC 2026 – All rights reserved
v
Technical Report ISO/IEC TR 5812:2026(en)
Information technology — Semantic metadata support in
office documents
1 Scope
This document outlines:
— descriptions of use cases for identified application domains that benefit from additional semantic
metadata support in office documents;
— descriptions of semantic metadata requirements from experts of identified application domains;
— specifications of possible technical strategies for associating domain-specific semantic metadata with
office documents entities.
This document applies to users of office documents, for example, publishers, librarians, archivists, data
scientists and other document users, as well as office software developers and service system integrators.
This document focuses on the methods of supporting various kinds of metadata, rather than on metadata
itself.
2 Normative references
There are no normative references in this document.
3 Terms and definitions
No terms and definitions are listed in this document.
ISO and IEC maintain terminology databases for use in standardization at the following addresses:
— ISO Online browsing platform: available at https:// www .iso .org/ obp
— IEC Electropedia: available at https:// www .electropedia .org/
4 Document semantic metadata
ISO/IEC 11179-1:2023, 5.1 describes metadata as “data that define and describe other data.”
ISO 19115-1:2014, 4.10 describes metadata as “information about a resource.” As for semantic metadata,
W3C describes metadata as enriching data with formal semantics, using ontologies or vocabularies.
[5]
Dublin Core Metadata Initiative (DCMI) describes metadata as being grounded in a formal knowledge
[6]
model, where terms are linked to defined concepts in a shared ontology. In the document, which parts
belong to semantic metadata and which do not depends on the requirements of the application. That is, if
the application requires machines to interpret only certain parts of the document, then those parts can be
represented as semantic metadata.
Metadata selection differs among communities, each with specific information requirements to be conveyed
by means of metadata. Before the advent of the information society, metadata were widely used for retrieving
books in libraries and identifying publications among publishers, with most of this metadata being manually
created. Currently, most of the metadata attached to various data sources come from diverse origins, some
of which are manually created or annotated, some inherited or translated from other documents and some
© ISO/IEC 2026 – All rights reserved
automatically generated by machines. Regardless of the method, metadata processing is time-consuming
and labour-intensive. Therefore, it is essential to establish a proper mechanism to record, save and utilize
metadata, facilitating its reuse and enhancing interoperability. However, in the field of document processing,
especially with editable office documents, current support for metadata is incomplete. Therefore, finding a
more effective way to support document semantic metadata is crucial. This is also the focus of this document.
5 Office document communities
Because different communities of document users have different needs and find different elements of office
documents relevant at different times, it is difficult to image "typical use cases" where additional document
semantic support is beneficial. However, it is possible to identify communities of office document users who
use documents in similar ways and are likely to find similar semantic metadata elements relevant to their
work. This document results from an investigation into how communities in publishing, machine learning
(see ISO/IEC 22989) (especially natural language processing, (NLP) (see ISO/IEC 23053), and both libraries
and archives utilize office documents. The aim was to understand which functions of office documents
are relevant to the usage patterns of members of these communities, thereby enabling the description of
document semantic support tailored to specific document uses.
Due to space limitations in this document, it is not feasible to enumerate all use cases utilized by various
communities. For instance, as collaborative editing technology has gained popularity, it requires the
inclusion of information from multiple authors at the paragraph level. Even if some paragraphs are authored
by non-humans (e.g. generative model, see ISO/IEC 23053), it is still necessary to include metadata to
indicate the source of the content.
6 Typical use cases
6.1 Themes
Semantic metadata added to a document aims to describe its content in an interoperable and machine-
readable way. Semantic metadata can be particularly useful for a variety of purposes, including and not
limited to information discovery and retrieval, document organization and presentation, document content
analysis and processing, and long-term document preservation.
Based on the foregoing, only a few representative communities and their common use cases are outlined.
Here, some typical use cases are grouped under the four themes (outlined in 6.1.1, 6.1.2, 6.1.3 and 6.1.4) and
discussed from the perspectives of the three groups of communities we introduced in Clause 5.
6.1.1 Information discovery and retrieval
Semantic metadata can be used to improve information discovery and retrieval in a number of ways. It can
help to improve the discoverability and accessibility of information within office documents, as it allows
users to search for and locate specific pieces of information within a document or a group of documents
more easily. There are a few examples of how semantic metadata can be used in different contexts.
From publishers’ view: Throughout the life cycle of a book or magazine, some major activities are supported
by semantic metadata. For example, metadata describing the publisher’s content enhances information
discovery and retrieval. A well-crafted description greatly influences the result of online searching for any
product. Publishers can use semantic metadata to describe the subject matter, authors and contributors,
publishing organization, publication date, title and other important document details. Additionally, they
can tag chapters and sections of a book with relevant keywords. This semantic metadata can then be used
by search engines, bookseller websites, libraries and other information systems to help users find relevant
documents or specific topics more easily.
In addition to supporting business requirements, semantic metadata also assists publishers in a role of
content provider to meet user and consumer requirements. Journalists, for example, use semantic metadata
to quickly locate relevant news events in manuscripts. The same manuscript can be tagged by different
newspapers in different ways, highlighting the need for OOXML/ODF format documents to support diverse
vocabularies from different sources. Researchers can use semantic metadata to annotate scientific papers
© ISO/IEC 2026 – All rights reserved
with specific techniques or technologies mentioned in the paper, simplifying the process for other researchers
to find and reference these papers. Similarly, businesses can use semantic metadata to tag documents with
names of projects, clients, or stakeholders, facilitating easier document discovery and retrieval.
From librarians’ view: Librarians and archive departments also use semantic metadata to describe the
collections in their libraries. For instance, a librarian can use semantic metadata to describe the author,
publisher, category and other important details about a book or journal article. This semantic metadata can
then be used by library catalogues, online databases and other information systems to help users locate
relevant materials more easily or to help archivists make classifications more precisely.
From machine learning practitioners’ view: Semantic metadata can also be used in the fields of machine
learning and artificial intelligence. For instance, a machine learning practitioner can use semantic metadata
to annotate a data set with information about the data, such as the types of objects or features presented.
This application of semantic metadata in office documents can aid machine learning practitioners to build
more accurate and efficient information retrieval systems. For example, they can use the metadata to filter
search results to only include documents authored by a specific individual or created within a certain time
frame. Additionally, semantic metadata can enhance the relevance of each document to a specific search
query, ensuring the most pertinent documents are prioritized in search results.
6.1.2 Document organization and presentation
Ideally, metadata are recorded and evolve throughout the lifecycle of a document rather than being recreated
and stored multiple times in multiple places. Thus, semantic metadata can benefit document organization
and presentation.
From publishers’ view: Publishers use identity metadata to provide unique marking for publications or
content resources, such as International Standard Book Number(ISBN), International Standard Serial
Number(ISSN), and International Standard Audio-Visual Product Code. In the network and digital era, the
role of identity metadata has been increasingly prominent, enabling the assignment of unique identification
symbols for information resources and their description with attached descriptive metadata. Publishers
also use semantic metadata to organize and present their documents in a more logical way. For instance,
a publisher can create a searchable index or table of contents for a book using semantic metadata, helping
readers to find specific chapters or sections more easily. Semantic metadata can also be used to group
related documents or to highlight the most important or relevant content. Moreover, publishers need to
retrieve the most recent version of documents or to verify and correct the typesetting pattern of these texts.
For example, formatting choices in the document sometimes imply some semantics which is not otherwise
explicitly expressed.
From librarians’ view: A librarian can use semantic metadata to group books or classify documents in a more
user-friendly library catalogue, describing details such as format, language, author, publisher and subject.
From the archivists’ view, an archivist needs to convert an office document into a usable PDF document, an
e-pub or a website. An office document, drafted in OOXML/ODF format using office software like Microsoft
Office or Open Office, is expected to have its semantic metadata converted into XMP format. This ensures
the semantic metadata to remain extractable from the office document, allowing it to be placed in the PDF
document or even enabling a round-trip transformation back to the original format.
From machine learning practitioners’ view: Machine learning practitioners can use semantic metadata to
train machine learning models to understand the meaning and context of text, to classify documents based
on their topics, or to extract structured data from unstructured text. Consequently, this allows information
to be presented in a more organized, meaningful, relevant and user-friendly manner.
6.1.3 Document content analysis and processing
From publishers’ view: Publishers of newspapers and periodicals can use semantic metadata to identify
trends or patterns in the content of their documents, or to discern the most popular topics or authors among
their readers. Media executives analyse which parts of articles their readers care about the most, and
whether their reviews are positive or negative. Probably they are also interested in determining which parts
of an article took the author the longest to write. For writers, there is a desire for writing tools not only to
automatically suggest the most valuable reference materials but also to support auto-drafting. Furthermore,
© ISO/IEC 2026 – All rights reserved
a writer expects the tool to automatically generate an after-book-index when compiling a textbook and allow
the machine to typeset the entire document automatically according to readers’ preferences.
From librarians’ view: Librarians can use semantic metadata to create customized views or presentations
of information, such as generating summaries of key points from a group of documents or creating filtered
lists of documents.
From machine learning practitioners’ view: Machine learning practitioners can use semantic metadata to
classify documents, identify relevant information and train machine learning models for more accurate
and efficient document classification, processing, analysing and outcome prediction. This improves the
interpretability of machine learning models. Existing metadata supports the NLP process, while metadata
generated by NLP enhances document retrievability and usability. Often, NLP techniques are employed to
search and retrieve documents due to a lack of sufficient semantic metadata in these documents. People
need to download, analyse and process the documents to acquire information. Using artificial intelligence
methods, intelligence analysis agencies can extract textual semantic information from a large number
of document samples (including web samples). This information includes identified names, locations,
organization names from different parts of documents such as columns, text and advertisements. These
extracted results need to be automatically tagged as semantic metadata. Furthermore, this semantic
metadata can record the results of NLP processing, enabling other users to directly utilize the previously
captured information, which eliminates the need for re-extraction when they first access these documents.
Metadata can also be artificially added or derived from different stages of document processing. Additionally,
OOXML/ODF format documents can support different namespaces and vocabularies from different sources.
A generative model (GM) is a model that generates new data instances resembling training data (e.g. text,
images, audio) (see ISO/IEC 23053). The core idea behind GM is to utilize artificial intelligence algorithms
to generate content. By training models and learning from large amounts of data, GM can generate content
related to specific prompts or guidance. For example, by inputting keywords, descriptions, or samples, GM
can generate matching articles, images, audio files and more. Currently, many document processing tools
[9]
have incorporated support for GM, such as Microsoft Office or WPS. These products allow computers to
automatically generate outlines, chapter contents, formulas in spreadsheets, abstracts, etc.; they also assist
in automatic document formatting (such as generating presentation slides) or layout optimization. While the
positive impact of GM has been widely recognized, its negative effects have also raised concerns. Compared
to previous technologies, GM brings greater security risks such as generating incorrect or inaccurate content,
altering the stylistic characteristics of writing, leading to privacy breaches or copyright infringement issues.
It is necessary to include specific metadata in documents that declare the source of training data for GMs
and inference results for better traceability purposes. This metadata includes identifiers and versions of
large-scale language model (LLM) (see ISO/IEC 23053) used in generation processes along with domain
classifications for applications employed; authorization information; generation time; digital signatures for
generated content, among others. Such metadata contribute to the security, interpretability, fairness and
privacy protection aspects of GM.
6.1.4 Long-term document preservation
From publishers’ view: Although publishers generally prioritize content reuse over long-term preservation
of documents, as seen in single publishing practices, semantic metadata can be a useful tool to help
reorganize the content of the book. This objectively ensures that the document has a long-term value and
prolongs the life of the document. The metadata that support the long-term preservation of resources
typically encompass detailed information regarding formats, production processes, protection measures,
data migration strategies, preservation responsibilities and more.
From librarians’ view: Semantic metadata preserve the meaning and significance of documents over time,
especially historical or archival documents where the context or background information is not immediately
apparent to future readers. Librarians save semantic metadata within the document to facilitate later
retrieval of relevant documents, rather than storing the entire document itself.
From machine learning practitioners’ view: In office documents, some users tend to tag semantic metadata
inside the document, allowing office software to recognize and process it. Others choose to store semantic
metadata outside the document, mapping it to specific fragments of the document so that the office document
format remains unchanged while recording semantic metadata.
© ISO/IEC 2026 – All rights reserved
6.2 General questions to consider
Based on the use cases discussed above, some general aspects of semantic metadata are concluded as
follows:
a) Relevance and accuracy: Semantic metadata are expected to be closely related to the content of the
document, accurately describing the document's subject, context and other relevant information
according to different scenarios, while avoiding errors or inconsistencies that can confuse or mislead
users. From a technical point of view, semantic metadata are explicitly related to the entire or parts of
the document. Authors of the document add or modify semantic metadata, which is then proofread by a
professional publishing house to ensure consistency with the edited document. Additionally, document
editing tools can set permissions to allow or disallow metadata editing. To facilitate the simultaneous
editing of documents and semantic metadata, the aim is to support semantic metadata with features
that are compatible with most editable document formats and processing tools. The annotation and
bookmark features of a document can be a good choice.
b) Interoperable and consistency: Semantic metadata are best expressed in an interoperable manner,
so that it can be easily interpreted and used by different users and systems. From a technical point
of view, the expression of semantic metadata adheres to widely adopted technical specifications to
enhance compatibility. For example, industry vocabularies expressed in XML Schema and knowledge
representation frameworks in Resource Description Framework (RDF) (see Clause 7) are based on
Subject-Predicate-Object. These standards are popular and effectively cover the metadata used by
various industries as discussed in the above use cases in 6.1.1, 6.1.2, 6.1.3 and 6.1.4.
c) Us
...



