• Yet another text processing application

    From anton@anton@mips.complang.tuwien.ac.at (Anton Ertl) to comp.lang.forth on Sun Sep 27 10:09:32 2026
    From Newsgroup: comp.lang.forth

    The task at hand is to add links to the HTML version of Gforth's NEWS
    file. The source code for the NEWS file is in markdown format; the
    information about the link for each word comes from the HTML version
    of the word index. You can find these files online:

    https://cgit.git.savannah.gnu.org/cgit/gforth.git/tree/NEWS.md https://net2o.de/gforth-1.0/Word-Index.html

    So a data-flow program of the files and tools involved is:

    NEWS.md --[pandoc]--> NEWS-nolinks.html -------\
    [news-links.fs] -> NEWS.html
    /
    ... -> gforth.texi --[texi2any]-> Word-Index.html

    The input files are text files, and the output file is a text file.
    Some people say that Forth is a poor fit for such tasks. I don't
    think so. Judge for yourself. The whole program is at

    https://cgit.git.savannah.gnu.org/cgit/gforth.git/tree/news-links.fs

    and I will show its pieces here. I don't think that it would have
    been much shorter in GNU awk (a language I use regularly for such
    tasks).

    The NEWS-nolinks.html input and NEWS.html output can be seen at:

    http://www.complang.tuwien.ac.at/anton/tmp/NEWS-nolinks.html http://www.complang.tuwien.ac.at/anton/tmp/NEWS.html

    So let's take a look at the code, which is unashamedly
    Gforth-specific.

    "https://net2o.de/gforth-1.0/" 2constant url-prefix

    Define this at the start, so we can change it easily when the base
    URL changes.

    "doc/gforth_html/Word-Index.html" slurp-file 2constant linkdata

    This reads (slurps) the whole file into memory and stores the
    resulting (long) string.

    wordlist constant links

    For every word in the index, this wordlist contains a link to the
    documentation.

    : parse-linkdata ( -- )
    case
    "<td class=\"printindex-index-entry\"><a href=\""
    string-parse nip 0= ?of endof
    "\"><code>" string-parse dup 0= ?of 2drop endof
    parse-name dup 0= ?of 2drop 2drop endof
    [: links set-current nextname 2constant ;] current-execute
    next-case ;

    This word expects the linkdata string in SOURCE, and Forth parsing
    words can work on it. STRING-PARSE uses a string rather than a
    single character as delimiter. The benefit over doing this with
    SEARCH is that the searched/parsed string does not have to be
    manipulated on the stack. The word index contains strings with the
    following pattern for the various words:

    <td class="printindex-index-entry"><a href="URL"><code>WORD ( ...

    We are interested in producing a lookup table where we can lookup
    the URL given the word. The lookup table is the wordlist LINKS.
    With the first STRING-PARSE we get to the start or URL, with the
    second we parse URL and get to the start of WORD. We parse word
    with PARSE-NAME and put a 2CONSTANT with the name WORD and the value
    URL in the wordlist LINKS. The CURRENT-EXECUTE saves and restores
    the current wordlist, so that the xt it executes is free to change
    it. Note that the use of STRING-PARSE and 2CONSTANT means that the
    URL string is inside LINKDATA.

    This word uses a loop with three exit conditions, implemented with
    CASE NEXT-CASE and, for the exit conditions, ?OF.

    linkdata `parse-linkdata execute-parsing

    Here EXECUTE-PARSING turns the string in LINKDATA into the input
    stream (such that parsing words work) while it executes the xt of
    PARSE-LINKDATA. `PARSE-LINKDATA is a state-independent way to get
    the xt of PARSE-LINKDATA.

    "NEWS-nolinks.html" slurp-file 2constant news

    The other input file is read completely and becomes NEWS.

    : parse-code ( -- )
    source drop >r begin ( R: c-addr )
    parse-name r@ third r> - type 2dup + >r
    dup while
    2dup links find-name-in dup if
    .\" <a href=\"" url-prefix type
    name>interpret execute type .\" \">"
    type ." </a>"
    else
    drop type
    then
    repeat
    rdrop 2drop ;

    In NEWS-nolinks.html certain substrings contain sequences of word
    names. This word parses one such sequence with PARSE-NAME, checks
    if the found word name has an entry in LINKS, and if so, outputs the
    word with an HTML link, otherwise without. The link gets the URL
    prefix defined earlier.

    : parse-news ( -- )
    case
    "<code>" string-parse dup 0= ?of 2drop endof
    type "<code>" type
    "</code>" string-parse dup 0= ?of 2drop endof
    `parse-code execute-parsing "</code>" type
    next-case ;

    The substrings containing the word names are surrounded by
    <code>...</code>, so this word searches (parses) for <code>, then
    for </code>, and the result of the latter parse (the string between)
    is passed to PARSE-CODE with `PARSE-CODE EXECUTE-PARSING. This word
    also has to output the result of the first PARSE-STRING and the
    strings parsed for.

    news `parse-news execute-parsing

    Let PARSE-NEWS work on NEWS.


    One shortcoming of the current version of this program is that it
    produces links for some words in the prose of the older NEWS. This
    could be mostly fixed by only putting links on capitalized words (this
    is complicted by cases where the HTML encoding of, e.g. ">" as "&gt;"
    contains lower-case characters). For now this is low on the ToDo
    list.

    This program uses the following words that are not standard and have
    not been accepted for standardization:

    slurp-file: used twice, quite useful.

    string-parse (4x) execute-parsing (3x): make it much easier to search
    through the files/strings that if we had used SEARCH. What is missing
    (also for search), but was not needed here is to search for one of
    several strings; Gforth has a regexp package that may be usable for
    that.

    ?of (5x) next-case (2x): Very useful for defining loops with several
    exits. VFX also has these two words. In the present cases the
    NEXT-CASE is followed by ";", so some will argue that using a BEGIN
    AGAIN loop with IF ... EXIT THEN instead of ?OF ... ENDOF would have
    worked just as well, and there is something to it. But sometimes I
    want to put a debugging tracer (~~) when a word is exited, and having
    several EXITs in addition to ";" means I have to insert several
    tracers. Using multiple WHILEs would have resulted in
    hard-to-understand code for getting the right stack for each WHILE.

    nextname: Often very useful for defining words where you do not have
    the names in parseable form. However, looking at the code again, I
    actually have the name in parseable form, and if we can live with the empty-parse checking that 2CONSTANT does (and that should be ok for
    this program), one can simplify PARSE-LINKDATA as follows:

    : parse-linkdata ( -- )
    case
    "<td class=\"printindex-index-entry\"><a href=\""
    string-parse nip 0= ?of endof
    "\"><code>" string-parse dup 0= ?of 2drop endof
    [: links set-current 2constant ;] current-execute
    next-case ;

    current-execute: wrappers for preserving certain state that is changed
    inside are useful and (if written correctly) prevent the problem that
    the programmer writes restoration code that does not work if the
    wrapped code throws.

    third: can also be written as 2 PICK.

    RDROP: can also be written as R> DROP.

    .\": can also be written as "<string>" TYPE.


    Maybe just as notable is what is not used: a string stack or even
    Gforths $tring words. I find that the cases where such features help
    in my work are few and far between.

    - anton
    --
    M. Anton Ertl http://www.complang.tuwien.ac.at/anton/home.html
    comp.lang.forth FAQs: http://www.complang.tuwien.ac.at/forth/faq/toc.html
    New standard: https://forth-standard.org/
    EuroForth 2026 CFP: http://www.euroforth.org/ef26/cfp.html
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From albert@albert@spenarnc.xs4all.nl to comp.lang.forth on Tue Sep 29 01:10:10 2026
    From Newsgroup: comp.lang.forth

    In article <2026Sep27.120932@mips.complang.tuwien.ac.at>,
    Anton Ertl <anton@mips.complang.tuwien.ac.at> wrote:
    <SNIP>
    The input files are text files, and the output file is a text file.
    Some people say that Forth is a poor fit for such tasks. I don't
    think so. Judge for yourself. The whole program is at

    Hear,hear!
    <SNIP>
    Maybe just as notable is what is not used: a string stack or even
    Gforths $tring words. I find that the cases where such features help
    in my work are few and far between.
    Hear,hear!

    In my opinion string stack are just an academic exercise.


    - anton

    Groetjes Albert
    --
    The Chinese government is satisfied with its military superiority over USA.
    The next 5 year plan has as primary goal to advance life expectancy
    over 80 years, like Western Europe.
    --- Synchronet 3.22a-Linux NewsLink 1.2
  • From Hans Bezemer@the.beez.speaks@gmail.com to comp.lang.forth on Tue Sep 29 15:27:18 2026
    From Newsgroup: comp.lang.forth

    On 29-09-2026 01:10, albert@spenarnc.xs4all.nl wrote:
    In my opinion string stack are just an academic exercise.

    In my opinion it is a vital part of 4tH's preprocessor -- which I praise
    every single day.

    Hans Bezemer

    --- Synchronet 3.22a-Linux NewsLink 1.2