• Re: Usenet-Rewind: search engine for Usenet messages back to 1981

    From Chris J Dixon@chris@cdixon.me.uk to uk.media.radio.archers on Sun Sep 6 10:10:36 2026
    From Newsgroup: uk.media.radio.archers

    Over in alt.folklore.computers

    Craig Stadler wrote:

    I spent the last year gathering and normalizing as many resources for
    Usenet messages and built a search engine over public Usenet
    discussions going back to 1981. Posting here because this seemed like
    the one group that would actually care.

    https://www.usenet-rewind.com/

    Since Google Groups stopped indexing new Usenet content a
    while back, and search over its historical archive has always been
    rough- I wanted something that treats this as a real research
    archive -- full-text search, date-range/filtering,etc and most
    importantly as much Usenet data as possible going back as far as I could >find.

    The index currently holds (to date) roughly 980 million messages and is
    still growing. Sources are archival backups, data donations, and ongoing >crawls of several thousand still-active public NNTP servers. Binary and >yEnc-encoded content is stripped where possible to keep the index
    focused on text.

    The index was built using Apache Solr 10, MariaDB 12 (rocksdb), Ubuntu >22.04.5 LTS & custom Python scripts, all on nvme disk.

    This is public archival content, same legal basis as DejaNews and
    Google Groups before it. There's a removal process for anyone who
    wants their own posts taken down, and author contact info is masked
    by default.

    I thought that some might find it interesting.

    Chris
    --
    Chris J Dixon Nottingham UK
    chris@cdixon.me.uk @ChrisJDixon1

    Plant amazing Acers.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From BrritSki@rtilbury@gmail.com to uk.media.radio.archers on Sun Sep 6 10:16:15 2026
    From Newsgroup: uk.media.radio.archers

    On 06/09/2026 10:10, Chris J Dixon wrote:
    Over in alt.folklore.computers

    Craig Stadler wrote:

    I spent the last year gathering and normalizing as many resources for
    Usenet messages and built a search engine over public Usenet
    discussions going back to 1981. Posting here because this seemed like
    the one group that would actually care.

    https://www.usenet-rewind.com/

    Since Google Groups stopped indexing new Usenet content a
    while back, and search over its historical archive has always been
    rough- I wanted something that treats this as a real research
    archive -- full-text search, date-range/filtering,etc and most
    importantly as much Usenet data as possible going back as far as I could
    find.

    The index currently holds (to date) roughly 980 million messages and is
    still growing. Sources are archival backups, data donations, and ongoing
    crawls of several thousand still-active public NNTP servers. Binary and
    yEnc-encoded content is stripped where possible to keep the index
    focused on text.

    The index was built using Apache Solr 10, MariaDB 12 (rocksdb), Ubuntu
    22.04.5 LTS & custom Python scripts, all on nvme disk.

    This is public archival content, same legal basis as DejaNews and
    Google Groups before it. There's a removal process for anyone who
    wants their own posts taken down, and author contact info is masked
    by default.

    I thought that some might find it interesting.


    Very. Thanks, will try later...

    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Nick Odell@nickodell49@yahoo.ca to uk.media.radio.archers on Sun Sep 6 11:24:20 2026
    From Newsgroup: uk.media.radio.archers

    On Sun, 06 Sep 2026 10:10:36 +0100, Chris J Dixon <chris@cdixon.me.uk>
    wrote:

    Over in alt.folklore.computers

    Craig Stadler wrote:

    I spent the last year gathering and normalizing as many resources for >>Usenet messages and built a search engine over public Usenet
    discussions going back to 1981. Posting here because this seemed like
    the one group that would actually care.

    https://www.usenet-rewind.com/

    Since Google Groups stopped indexing new Usenet content a
    while back, and search over its historical archive has always been
    rough- I wanted something that treats this as a real research
    archive -- full-text search, date-range/filtering,etc and most
    importantly as much Usenet data as possible going back as far as I could >>find.

    The index currently holds (to date) roughly 980 million messages and is >>still growing. Sources are archival backups, data donations, and ongoing >>crawls of several thousand still-active public NNTP servers. Binary and >>yEnc-encoded content is stripped where possible to keep the index
    focused on text.

    The index was built using Apache Solr 10, MariaDB 12 (rocksdb), Ubuntu >>22.04.5 LTS & custom Python scripts, all on nvme disk.

    This is public archival content, same legal basis as DejaNews and
    Google Groups before it. There's a removal process for anyone who
    wants their own posts taken down, and author contact info is masked
    by default.

    I thought that some might find it interesting.

    Yes indeed! Thank you for posting that.

    I've saved the link - although I suppose I could always find it again
    on Usenet-Rewind :-)

    Nick
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Chris J Dixon@chris@cdixon.me.uk to uk.media.radio.archers on Sun Sep 6 13:17:19 2026
    From Newsgroup: uk.media.radio.archers

    Nick Odell wrote:

    On Sun, 06 Sep 2026 10:10:36 +0100, Chris J Dixon <chris@cdixon.me.uk>
    wrote:

    Over in alt.folklore.computers

    Craig Stadler wrote:

    I spent the last year gathering and normalizing as many resources for >>>Usenet messages and built a search engine over public Usenet
    discussions going back to 1981. Posting here because this seemed like
    the one group that would actually care.

    https://www.usenet-rewind.com/

    Since Google Groups stopped indexing new Usenet content a
    while back, and search over its historical archive has always been
    rough- I wanted something that treats this as a real research
    archive -- full-text search, date-range/filtering,etc and most >>>importantly as much Usenet data as possible going back as far as I could >>>find.

    The index currently holds (to date) roughly 980 million messages and is >>>still growing. Sources are archival backups, data donations, and ongoing >>>crawls of several thousand still-active public NNTP servers. Binary and >>>yEnc-encoded content is stripped where possible to keep the index
    focused on text.

    The index was built using Apache Solr 10, MariaDB 12 (rocksdb), Ubuntu >>>22.04.5 LTS & custom Python scripts, all on nvme disk.

    This is public archival content, same legal basis as DejaNews and
    Google Groups before it. There's a removal process for anyone who
    wants their own posts taken down, and author contact info is masked
    by default.

    I thought that some might find it interesting.

    Yes indeed! Thank you for posting that.

    I've saved the link - although I suppose I could always find it again
    on Usenet-Rewind :-)

    Error! Recursion.

    Chris
    --
    Chris J Dixon Nottingham
    '48/33 M B+ G++ A L(-) I S-- CH0(--)(p) Ar- T+ H0 ?Q
    chris@cdixon.me.uk @ChrisJDixon1
    Plant amazing Acers.
    --- Synchronet 3.22a-Linux NewsLink 1.2