What Is XML and What Is It Used For?
XML (Extensible Markup Language) is a text-based format for storing and transporting structured data. You use it when two systems need to exchange information in a way both can parse, when a document needs to be both machine-readable and human-readable, or when you need a strict, self-describing data format. It is not a programming language and does not do anything on its own — it only describes data. If you just need to display a page in a browser, you want HTML, not XML.
The core idea: you invent the tags
HTML comes with a fixed set of tags (<p>, <h1>, <div>) that mean specific things to a browser. XML has no predefined tags at all. You define your own names to describe your data:
<book isbn="978-0-13-110362-7">
<title>The C Programming Language</title>
<author>Brian Kernighan</author>
<author>Dennis Ritchie</author>
<year>1988</year>
</book>
Nothing here is a built-in XML keyword. book, title, and author are names the author chose. That is what "extensible" means: the vocabulary is open-ended, and the meaning comes from whatever agreement exists between the sender and receiver.
Basic syntax rules
XML is strict. A parser will reject a document that breaks these rules rather than guess what you meant:
- One root element. Every document has exactly one outermost element that contains everything else.
- Tags must close and nest properly.
<a><b></b></a>is valid;<a><b></a></b>is not. - Case matters.
<Title>and<title>are different elements. - Attributes are quoted.
<book isbn="123">works;<book isbn=123>does not. - Special characters are escaped. Use
<for<,&for&, and so on.
A well-formed document follows these rules. A valid document additionally conforms to a schema (DTD or XSD) that defines which elements and attributes are allowed. Well-formedness is required; validity is optional and only matters when you have agreed on a schema.
XML vs. HTML
They share a syntax heritage but solve different problems.
| XML | HTML | |
|---|---|---|
| Purpose | Describe and transport data | Display content in a browser |
| Tags | You define them | Fixed, predefined set |
| Error handling | Strict — stops on malformed input | Lenient — browsers recover and render anyway |
| Case sensitivity | Case-sensitive | Case-insensitive |
| Who consumes it | Applications, parsers | Browsers, humans |
The practical consequence: a single missing closing tag can break an entire XML pipeline, while the same mistake in HTML usually just renders slightly wrong. That strictness is a feature when data integrity matters.
Common uses
- Sitemaps. A sitemap is an XML file listing a site's URLs so search engines can crawl them. The XML Sitemaps Generator, for example, produces sitemaps in this format and can submit them to Google, Bing, and others. A minimal entry looks like
<url><loc>https://example.com/</loc></url>, and generators can automatically set a last-modified date and a priority value between 0.0 and 1.0 based on page depth. - RSS and Atom feeds. Syndication formats are XML. Each item is an element with a title, link, and publication date.
- Configuration files. Many tools (Maven, Android layouts, some enterprise software) read settings from XML.
- Data exchange between systems. Especially in older enterprise, finance, and government integrations where a formal schema is required.
- Office and document formats.
.docx,.xlsx, and.svgare ZIP archives full of XML files.
XML vs. JSON: which to choose
JSON has largely replaced XML for new web APIs because it is lighter and maps directly onto common data structures. The choice usually comes down to your constraints:
- Choose JSON for new REST APIs, browser-to-server communication, and config files where humans edit by hand. It is less verbose and natively understood by JavaScript.
- Choose XML when you need a formal schema for validation, when you must mix text and markup in the same document (as in documents and SVG), when you need comments and namespaces, or when you are integrating with a system that already speaks XML.
Neither is "better" in the abstract. If the other side of the conversation expects XML, send XML.
Reading and writing XML in practice
To read XML programmatically, use a parser rather than string manipulation — regular expressions break on nesting and escaping. Standard libraries exist in every major language (for example xml.etree.ElementTree in Python, DOMParser in JavaScript, SimpleXML in PHP). To write it, generate the structure through a library or serializer so that escaping and closing tags are handled for you.
A quick way to check a document by hand: open it in a browser or an XML validator. If it renders as a collapsible tree, it is well-formed; if you get a parse error, the message usually points to the line and the unclosed or mismatched tag.