How Webarchiv.cz Archives and Preserves Czech Websites
Webarchiv.cz is the Czech National Library's web archive, and it preserves Czech websites by harvesting them, storing the captured copies in a long-term digital repository, and making them accessible to the public through its browsing interface. It is relevant to you if you want to find a Czech website that has disappeared or changed, or if you run a Czech site and want it preserved. The archive is selective rather than exhaustive: it works both through contracts with publishers and through its own selection of sites, so not every Czech website is included.
What Webarchiv.cz is
Webarchiv.cz describes itself as "the Museum of Czech web." Its stated role is to collect web resources, archive them, and ensure long-term access to them. It is part of the National Library of the Czech Republic (NL CR).
Two things follow from that framing:
- The archive treats websites as cultural heritage, not just as temporary files.
- Its purpose is long-term access, so the value of a site in the archive grows as the original goes offline.
How archiving works
The archive does not host the live website. It captures a copy of the site at a point in time and stores that copy. The site's publisher is not required to keep the original online for the archived copy to remain reachable.
The process, as reflected on the site, has two main entry paths:
- Contracted websites. Webarchiv.cz concludes agreements with publishers to archive their sites. As of the site's current figures, it has concluded 4,769 contracts with publishers. Recent examples listed include Spirituality Studies Revue, Turistické informační centrum Vyškov, Nový Vyškov, Magazín Živá univerzita, and Centrum infekčních nemocí zvířat.
- Selected websites. Beyond contracts, the archive selects sites on its own initiative. The site lists a "Selection of contracted websites" with a link to the full list, and offers a "nominate a site" option, which means public suggestions feed into what gets archived.
The first website was harvested on 3 September 2001, so the collection spans more than two decades of the Czech web.
Scale of the collection
The archive reports holding 740 TB of data as of 23 September 2026. That figure is a useful signal of scope: this is a large, actively growing collection rather than a small sample.
| Item | Value |
|---|---|
| First harvest | 3 September 2001 |
| Data held | 740 TB (as of 23 September 2026) |
| Publisher contracts | 4,769 |
| Operator | National Library of the Czech Republic |
Why small local websites matter here
Webarchiv.cz gives particular attention to smaller local websites, on the reasoning that they are often the first to disappear from the internet. Local history, traditions, associations, and cultural events are increasingly documented only online, so when a small site goes down, that record can vanish with it.
Two examples from the archive's own selection show what this looks like in practice:
- Stará Jihlava — a site about the history of the town of Jihlava, now findable only in the archive. It contains historical photographs, articles on the town's history, and old maps. Publisher: Oubrecht, Lukáš.
- Historie obce Sobíňov — a detailed history of the village of Sobíňov in the Vysočina region, online until around 2017. It includes historical photographs and describes the village's parts, buildings, landmarks, and inhabitants. Publisher: Jágr, Miroslav.
Both are exactly the kind of material that would be hard to reconstruct if the original sites had not been captured.
How the public can use it
The archive is built for public access, not just preservation. From the site's navigation you can:
- Browse the collection.
- Explore topic collections.
- Search for a site or topic.
- Nominate a site you think should be archived.
The interface is available in Czech and English (CZ / EN), which makes it usable if you do not read Czech.
A practical way to use it: if you are researching a Czech town, village, association, or local event and the original website is gone or has been redesigned, search the archive for the site name or topic. If the site was captured, you can view the archived version instead of the dead link.
What to keep in mind
- Coverage is selective. Contracts and curation mean the archive does not contain every Czech website. A missing site does not necessarily mean it was rejected — it may simply never have been captured.
- You are viewing a snapshot. An archived copy reflects the site at the time of harvesting, not its current state.
- Nomination is an option. If a site you value is not in the archive, the "nominate a site" route is the documented way to flag it.
- The archive is run by a national institution. Its preservation mission is tied to the National Library of the Czech Republic, which is why long-term access, rather than short-term convenience, is the design goal.