I have a few hundered feet of bookshelves. I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are
readable and complete I'll dispose of the physical copy.
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
I've tried converting the background colour to white, and the scan postprocess tools available in the likes of Acrobat, but these always interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
I have a few hundered feet of bookshelves. I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are
readable and complete I'll dispose of the physical copy.
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
I've tried converting the background colour to white, and the scan postprocess tools available in the likes of Acrobat, but these always interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
"JM" <sunaecoNoChoppedPork@gmail.com> wrote in message news:umu38l9552gkf9fgeprn2r9k8el08rmjj4@4ax.com...
I have a few hundered feet of bookshelves. I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are
readable and complete I'll dispose of the physical copy.
I find it hard to get rid of physical copies but it's probably true that I should get rid of more than I do.
I have some but not many books which can't be found online as pdf or some other format.
I find it just as easy to lose the pdf file as the physical book. It has to be in
a folder somewhere.
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
Sounds like this may be because it's having to ocr images.
I've noticed that pdfs vary a lot in quality and size of file.
Bigger isn't necessarily better.
There are a lot of online tools to process PDFs. But, that becomes tedious when you have gigabytes of individual files that have to be "uploaded", processed and downloaded. And, you likely need to review each to ensure something hasn't been excised from the original that may have value.I've tried converting the background colour to white, and the scan
postprocess tools available in the likes of Acrobat, but these always
interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
https://www.pdfgear.com/
may or may not help.
On 8/16/2026 11:47 AM, Edward Rawde wrote:
"JM" <sunaecoNoChoppedPork@gmail.com> wrote in message news:umu38l9552gkf9fgeprn2r9k8el08rmjj4@4ax.com...
I have a few hundered feet of bookshelves. I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are
readable and complete I'll dispose of the physical copy.
I find it hard to get rid of physical copies but it's probably true that I >> should get rid of more than I do.
I used to *cling* to "books" as if they were sacred. No marks, no dog-ears, etc.
But, after a lifetime, you end up with *so* many that they start to make demands on you -- where they are stored, how htey are stored, how they
are organized, etc.
When I got my first Nook, I quickly realized that *all* of my novels
were just taking up space -- instead of putting them on microSD cards
and carrying the whole library with me (sadly, ereaders don't have
an effective way of organizing large collections -- they seem to be
tailored to folks reading the most recent crop of "Best Sellers")
Text books are more difficult as they often have illustrations and
other presentation aspects that are key to understanding the content.
Ditto research papers, data books, etc.
Databooks can often be DLed so that avoids that issue (except for
older items that have particular interest).
Research papers are almost always available in PDFs -- even if they
happen to be in 8.5x11 format.
Text books are the hardest to find, create and review.
I now use an 18" tabletPC for the items that can't easily be
"reflowed" like an epub. It's a bit large/heavy -- but no worse
than a similarly sized textbook.
I have some but not many books which can't be found online as pdf or some
other format.
I find it just as easy to lose the pdf file as the physical book. It has to be in
a folder somewhere.
As with every "file" that I have, I have a database that tracks its location (which medium, which folder/container, etc). Along with any duplicates that may be replicated elsewhere (e.g., projects often reference the same datasheets
so those naturally appear in each "project folder")
My database lets me annotate each entry with a freeform note (to myself). Sorting based on the *presence* of such a note is often enough to find what
I am interested in as its presence means the item was of sufficient interest for me to have made that annotation.
Or, look at the enclosing folder hierarchy ("Ah, this was part of Project X" and rely on personal memory to augment that structure)
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
Sounds like this may be because it's having to ocr images.
I've noticed that pdfs vary a lot in quality and size of file.
Bigger isn't necessarily better.
PDFs are just containers. You can put images, text, binaries, multimedia, etc.
into a PDF. How it *appears* is up to the person who crafts the PDF.
There are a lot of online tools to process PDFs. But, that becomes tedious when you have gigabytes of individual files that haveI've tried converting the background colour to white, and the scan
postprocess tools available in the likes of Acrobat, but these always
interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
https://www.pdfgear.com/
may or may not help.
to be "uploaded",
processed and downloaded. And, you likely need to review each to ensure something hasn't been excised from the original that may have value.
I have a few hundered feet of bookshelves. I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are
readable and complete I'll dispose of the physical copy.
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
I've tried converting the background colour to white, and the scan postprocess tools available in the likes of Acrobat, but these always interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
As with every "file" that I have, I have a database that tracks its location >> (which medium, which folder/container, etc). Along with any duplicates that >> may be replicated elsewhere (e.g., projects often reference the same datasheets
so those naturally appear in each "project folder")
I use databases too, but there's no way I could use a database to track the location
of a book because the database would never be up to date with the book's current location.
On 16/08/2026 19:13, JM wrote:
I have a few hundered feet of bookshelves.-a I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu).-a If they are
readable and complete I'll dispose of the physical copy.
I trust original paper copies of books to survive.
AI companies are buying up secondhand books at a prodigious rate and destroying
them to train AI ready for the singularity. Although what use "How to pass your
driving test in Wales" is to something intending world domination is a mystery
to me.
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
Image to text can work but you need to carefully proofread every page.
many books are already available in online digital libraries.
I've tried converting the background colour to white, and the scan
postprocess tools available in the likes of Acrobat, but these always
interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
The trick is to use a greedy histogram equalisation down to 4 colours which leaves enough margin to smooth the text and also lose the foxing. If they are
full colour scans then split to RGB first and use only the red channel (which
is usually pristine).
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
Smart histogram equalisation is the only way to do it.
Possibly with a bit of unsharp masking first to sharpen up the edges.
I trust original paper copies of books to survive. || |
On 8/16/2026 12:42 PM, Martin Brown wrote:E.g., this is one of my most precious books:
On 16/08/2026 19:13, JM wrote:
I have a few hundered feet of bookshelves.-a I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu).-a If they are
readable and complete I'll dispose of the physical copy.
I trust original paper copies of books to survive.
I'm not that confident.-a I have many paper books that have pages that
have become very brittle.-a Others where the pages "snapped off" at
the binding edge.
On 16/08/2026 19:13, JM wrote:
I have a few hundered feet of bookshelves. I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu). If they are
readable and complete I'll dispose of the physical copy.
I trust original paper copies of books to survive.
AI companies are buying up secondhand books at a prodigious rate and destroying them to train AI ready for the singularity. Although what use "How to pass your driving test in Wales" is to something intending world domination is a mystery to me.
But the PDF's in particular are often very slow to nagivate because of
the colouring and texture of thepage background, which can take quite
a while for the PDF program to generate.
Image to text can work but you need to carefully proofread every page.
many books are already available in online digital libraries.
I've tried converting the background colour to white, and the scan
postprocess tools available in the likes of Acrobat, but these always
interfere with the text (make it blurry) and mahe the PDF's unpleasant
to read.
The trick is to use a greedy histogram equalisation down to 4 colours
which leaves enough margin to smooth the text and also lose the foxing.
If they are full colour scans then split to RGB first and use only the
red channel (which is usually pristine).
Has anyone been through this and found a decent workflow to lighten
PDF backrounds without touching the text?
Smart histogram equalisation is the only way to do it.
Possibly with a bit of unsharp masking first to sharpen up the edges.
On 8/16/2026 12:42 PM, Martin Brown wrote:
On 16/08/2026 19:13, JM wrote:
I have a few hundered feet of bookshelves.-a I've been going through
the books and downloading scans of them, if I can find them somewhere.
These are mostly in PDF format (or sometimes djvu).-a If they are
readable and complete I'll dispose of the physical copy.
I trust original paper copies of books to survive.
I'm not that confident.-a I have many paper books that have pages that
have become very brittle.-a Others where the pages "snapped off" at
the binding edge.
Image to text can work but you need to carefully proofread every page.
many books are already available in online digital libraries.
You can store the image (monochrome, greyscale, color, etc.) AND the
OCR'd text "on top of it" (invisibly) in a PDF.-a This increases the
size of the PDF beyond what a "pure" PDF would require.-a But, is invaluable if you suspect the OCR process of being faulty.
You can also repeat the OCR with the imagery in the PDF as technology improves.
(But, this means you want to scan -- and preserve -- at high enough resolution
to enable that "post processing")
| Sysop: | Amessyroom |
|---|---|
| Location: | Fayetteville, NC |
| Users: | 74 |
| Nodes: | 6 (0 / 6) |
| Uptime: | 52:42:47 |
| Calls: | 1,101 |
| Calls today: | 1 |
| Files: | 1,339 |
| Messages: | 276,187 |