EPUBCheck is the reference validator, and retailers build their intake systems on it. Passing does not guarantee acceptance, since a book can be refused for things a validator never examines. But failing is the usual way a file is refused before anyone sees it.
Its output is accurate but hard to read: a code, a file, a line and column, and a sentence for someone who already knows the standard. This page explains it. A valid EPUB was broken fifteen ways and the responses recorded; the messages below are the ones you will see.
The packaging errors
An EPUB is a ZIP with rules about how the ZIP itself is built. They are checked first and broken most often, because zipping a folder in Finder or Explorer breaks them.
PKG-007 The mimetype file is compressed
The first file must be called mimetype, contain only that
string with no trailing newline, and be stored uncompressed. A
normal zip compresses everything, so re-zipping an unpacked EPUB by hand
almost always causes this.
PKG-006 The mimetype file is not first
Order inside a ZIP is not alphabetical, and file managers do not let you set it. The mimetype entry must come first, so a reading device can identify the format from the opening bytes.
The two numbers are the reverse of what people assume: compression is PKG-007, order is PKG-006. Read the message, not the number.
The reference errors
All the same idea: something points at something that is not there, usually a file that was renamed, moved or never included. These errors are common.
RSC-001 A file named in the manifest does not exist
The package promises a file the archive does not have, usually an image placed in the manuscript from your disk and never copied into the ebook.
RSC-007 A link or stylesheet points at a missing file
The same failure from the other end: a chapter links to a page that is not
there. It also catches a @font-face rule whose font file was
left behind, common when a theme is copied between projects.
RSC-012 A link points at an anchor that does not exist
The file is right and the #anchor after it is wrong. Often this is a footnote or endnote whose reference survived
an edit when the note did not.
RSC-011 The contents list points at something outside the reading order
The contents list links to a chapter that exists but is not in the sequence a reader moves through. Take this one seriously; see the last section of this page.
OPF-049 The reading order names a manifest entry that isn't there
This arrives with an RSC-005 saying the same thing in specification language: one problem, two messages. EPUBCheck's error count is not a count of problems.
RSC-006 and OPF-014 An image is loaded from the internet
The property "remote-resources" should be declared in the OPF file.
An ebook must contain everything it shows. An <img> with an
https:// source vanishes when the reader is offline, so it is
refused.
The document errors
RSC-016 — fatal — The XHTML is not well formed
XHTML is XML, and XML tolerates no mistakes: one unclosed tag stops the parse. It is reported twice, as FATAL RSC-016 and ERROR RSC-005, and nothing after the broken tag is checked. Fix it and run again before reading the rest.
Browsers repair broken HTML silently, which is why a file can look perfect in a browser and be fatally invalid here.
RSC-005 Something required is missing or in the wrong place
Error while parsing file: Exactly one "toc" nav element must be present
RSC-005 is the general “this does not match the schema” message, and the useful part is the sentence after the colon. The two above are common: a package with no title, and a navigation document whose contents list is not marked as the contents list.
OPF-003 — warning — A file is in the archive but not declared
Usually harmless: a leftover file, a .DS_Store, an editor's
backup. Clear it anyway, because readers have to download that dead weight.
Two things that pass but should not reassure you
Both produced 0 errors, and neither file is correct.
-
An obsolete DOCTYPE. The old XHTML 1.1 doctype instead of
<!DOCTYPE html>passed cleanly. Advice calling it an error is out of date. -
A language tag that is not a language tag.
englishinstead ofenpassed without comment. Reading devices use that field for hyphenation and text-to-speech, so a wrong value quietly degrades the book.
The failure EPUBCheck cannot catch
A whole chapter was removed from the reading order and from the contents list, with its file left in the archive. The result:
Messages: 0 fatals / 0 errors / 0 warnings / 0 infos
A valid ebook, missing a chapter. The standard only requires the parts an EPUB declares to agree with each other, and they did. Remove the chapter from one of the two places and you get RSC-011; remove it from both and it is a shorter book.
A clean EPUBCheck report means the file is well formed, not that it is your book. Only opening the book and counting the chapters catches a missing one.
Reading a report
- Fix every FATAL first. Parsing stopped there, so the rest of the report was written from an incomplete picture.
- One problem can produce several messages. Fix causes and run again, rather than working down the list.
- For RSC-005, read the sentence, not the code.
- Warnings are optional and cheap to fix. Clear them, so the next new one stands out.
Torcul validates with EPUBCheck itself. This page quotes the messages because paraphrasing them loses the part that tells you what to do.