On Sat, 26 Dec 2009 22:13:53 +0800 Paul Wise <[email protected]> wrote:
> Package: gpdftext > Severity: wishlist > > Some PDFs have a table of contents page that contains lines consisting > of the section title and then a series of periods (or other > characters) and then a page number. It would be nice if gpdftext > could detect these and replace the series of dots with just enough to > make the text fit on one line or just a space. An example of a > document with the kind of TOC I'm talking about is available from the > following URL: > > https://docs.indymedia.org/pub/Global/ImcEssayCollection/imc_future-v0.2.pdf In contrast, the kind of PDF's more commonly used in gpdftext are such as the ones available at: http://www.feedbooks.com/book/97.pdf 1. If you can offer a unique regular expression, it could work but the 'page number' regular expression tries to do something similar and often fails to match due to the difficulty of making a suitable reg exp that doesn't remove useful content, particularly from novels. 2. A lot of technical PDF's would have been generated from DocBook or similar, is a text version already available? - gpdftext has known problems with tabular data (due to limitations in the underlying poppler support). e.g. it cannot extract text from a PDF of a test CV. 3. If this was to be supported, someone's going to want the TOC links to be usable again once the text is saved as a new PDF. I can't see gpdftext gaining sufficient functionality to reassemble the TOC data myself. I'm not sure whether gpdftext can support what you are requesting. -- Neil Williams ============= http://www.data-freedom.org/ http://www.linux.codehelp.co.uk/ http://e-mail.is-not-s.ms/
pgp2ScNmZTPcD.pgp
Description: PGP signature

