On Sat, 26 Dec 2009 22:13:53 +0800
Paul Wise <[email protected]> wrote:

> Package: gpdftext
> Severity: wishlist
> 
> Some PDFs have a table of contents page that contains lines consisting
> of the section title and then a series of periods (or other
> characters) and then a page number. It would be nice if gpdftext
> could detect these and replace the series of dots with just enough to
> make the text fit on one line or just a space. An example of a
> document with the kind of TOC I'm talking about is available from the
> following URL:
> 
> https://docs.indymedia.org/pub/Global/ImcEssayCollection/imc_future-v0.2.pdf

In contrast, the kind of PDF's more commonly used in gpdftext are
such as the ones available at:
http://www.feedbooks.com/book/97.pdf

1. If you can offer a unique regular expression, it could work but the
'page number' regular expression tries to do something similar and
often fails to match due to the difficulty of making a suitable reg exp
that doesn't remove useful content, particularly from novels.

2. A lot of technical PDF's would have been generated from DocBook or
similar, is a text version already available? - gpdftext has known
problems with tabular data (due to limitations in the underlying poppler
support). e.g. it cannot extract text from a PDF of a test CV.

3. If this was to be supported, someone's going to want the TOC links
to be usable again once the text is saved as a new PDF. I can't see
gpdftext gaining sufficient functionality to reassemble the TOC data
myself.

I'm not sure whether gpdftext can support what you are requesting.

-- 


Neil Williams
=============
http://www.data-freedom.org/
http://www.linux.codehelp.co.uk/
http://e-mail.is-not-s.ms/

Attachment: pgp2ScNmZTPcD.pgp
Description: PGP signature

Reply via email to