Copyright (C) 2013-2018  Stephan Kreutzer

This file is part of html2epub1.

html2epub1 is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License version 3 or any
later version, as published by the Free Software Foundation.

html2epub1 is distributed in the hope that it will be useful,
but WITHOUT ANY WARRANTY; without even the implied warranty of
MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the
GNU Affero General Public License 3 for more details.

You should have received a copy of the GNU Affero General Public License 3
along with html2epub1. If not, see <http://www.gnu.org/licenses/>.



Description
-----------

html2epub1 is a tool to convert valid XHTML 1.0 Strict files to EPUB2. It is provided by https://gitlab.com/publishing-systems/automated_digital_publishing/.


Requirements
------------

A proper Java SDK must be installed to produce the *.class files. The source code is at least compatible with Java 1.6 and 1.7 (OpenJDK).


Build
-----

Type

    make

in the directory containing the package's source code.


Execution
---------

Type

    java html2epub1 config-file

in the html2epub1 directory, or anywhere on the shell if the path to the html2epub1 directory is appended to the CLASSPATH environment variable.


Usage
-----

The transformation process is driven by the XML-based configuration file, which needs to be passed as argument for the program call. The html2epub1 software package contains an example 'config.xml', where two input files are specified by <inFile>. Since every input file will reflect a "part" in the resulting EPUB file, the 'title' attribute specifies the name of the part for the table of contents. <xhtmlSchemaValidation> should be always set to 'true', for which to take effect

    http://www.w3.org/2002/08/xhtml/xhtml1-strict.xsd
    http://www.w3.org/2001/xml.xsd
    http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd
    http://www.w3.org/TR/xhtml1/DTD/xhtml-lat1.ent
    http://www.w3.org/TR/xhtml1/DTD/xhtml-symbol.ent
    http://www.w3.org/TR/xhtml1/DTD/xhtml-special.ent

should be downloaded and saved in the html2epub1 directory. Some of those files lack the legal permission to be redistributable by anyone, so those files need to be obtained directly from the W3C. Please note that the W3C artificially slows down the access to those files, blocks certain IPs, blocks certain browsers/user agents from accessing those files. If you view those files with a browser, an associated stylesheet may render the files into HTML output, so you need to view the source code of the page presented by the browser in order to successfully download the original resource. See

    http://www.w3.org/blog/systeam/2008/02/08/w3c_s_excessive_dtd_traffic/
    
for more details on this issue. <xhtmlSchemaValidation> might only be set to 'false', if the input XHTML files got already validated and weren't changed since. <outDirectory> specifies where html2epub1 will place the temporary output files. Those files can also be packed to EPUB manually by

    zip -X0 "out.epub" "mimetype"
    zip -Xr "out.epub" "META-INF" "content.opf" "toc.ncx" "page_1.xhtml" "page_2.xhtml" "image_1.png" "image_2.png"

German speakers might look at this short demonstration video: http://vimeo.com/89003773.


Missing features
----------------

@font-face file references in CSS, absolute paths for links and referenced images, table of contents based on <h1>, <h2>, <h3>, <h4>, <h5> and <h6> headers (with option to choose up to which level of headings), refuse to write invalid output (check the existence of anchor target IDs).


Contact
-------

See https://gitlab.com/publishing-systems/automated_digital_publishing/ or the contact form at the bottom of the page https://skreutzer.de/about.html.

