From Newsgroup: comp.lang.forth
The task at hand is to add links to the HTML version of Gforth's NEWS
file. The source code for the NEWS file is in markdown format; the
information about the link for each word comes from the HTML version
of the word index. You can find these files online:
https://cgit.git.savannah.gnu.org/cgit/gforth.git/tree/NEWS.md https://net2o.de/gforth-1.0/Word-Index.html
So a data-flow program of the files and tools involved is:
NEWS.md --[pandoc]--> NEWS-nolinks.html -------\
[news-links.fs] -> NEWS.html
/
... -> gforth.texi --[texi2any]-> Word-Index.html
The input files are text files, and the output file is a text file.
Some people say that Forth is a poor fit for such tasks. I don't
think so. Judge for yourself. The whole program is at
https://cgit.git.savannah.gnu.org/cgit/gforth.git/tree/news-links.fs
and I will show its pieces here. I don't think that it would have
been much shorter in GNU awk (a language I use regularly for such
tasks).
The NEWS-nolinks.html input and NEWS.html output can be seen at:
http://www.complang.tuwien.ac.at/anton/tmp/NEWS-nolinks.html http://www.complang.tuwien.ac.at/anton/tmp/NEWS.html
So let's take a look at the code, which is unashamedly
Gforth-specific.
"
https://net2o.de/gforth-1.0/" 2constant url-prefix
Define this at the start, so we can change it easily when the base
URL changes.
"doc/gforth_html/Word-Index.html" slurp-file 2constant linkdata
This reads (slurps) the whole file into memory and stores the
resulting (long) string.
wordlist constant links
For every word in the index, this wordlist contains a link to the
documentation.
: parse-linkdata ( -- )
case
"<td class=\"printindex-index-entry\"><a href=\""
string-parse nip 0= ?of endof
"\"><code>" string-parse dup 0= ?of 2drop endof
parse-name dup 0= ?of 2drop 2drop endof
[: links set-current nextname 2constant ;] current-execute
next-case ;
This word expects the linkdata string in SOURCE, and Forth parsing
words can work on it. STRING-PARSE uses a string rather than a
single character as delimiter. The benefit over doing this with
SEARCH is that the searched/parsed string does not have to be
manipulated on the stack. The word index contains strings with the
following pattern for the various words:
<td class="printindex-index-entry"><a href="URL"><code>WORD ( ...
We are interested in producing a lookup table where we can lookup
the URL given the word. The lookup table is the wordlist LINKS.
With the first STRING-PARSE we get to the start or URL, with the
second we parse URL and get to the start of WORD. We parse word
with PARSE-NAME and put a 2CONSTANT with the name WORD and the value
URL in the wordlist LINKS. The CURRENT-EXECUTE saves and restores
the current wordlist, so that the xt it executes is free to change
it. Note that the use of STRING-PARSE and 2CONSTANT means that the
URL string is inside LINKDATA.
This word uses a loop with three exit conditions, implemented with
CASE NEXT-CASE and, for the exit conditions, ?OF.
linkdata `parse-linkdata execute-parsing
Here EXECUTE-PARSING turns the string in LINKDATA into the input
stream (such that parsing words work) while it executes the xt of
PARSE-LINKDATA. `PARSE-LINKDATA is a state-independent way to get
the xt of PARSE-LINKDATA.
"NEWS-nolinks.html" slurp-file 2constant news
The other input file is read completely and becomes NEWS.
: parse-code ( -- )
source drop >r begin ( R: c-addr )
parse-name r@ third r> - type 2dup + >r
dup while
2dup links find-name-in dup if
.\" <a href=\"" url-prefix type
name>interpret execute type .\" \">"
type ." </a>"
else
drop type
then
repeat
rdrop 2drop ;
In NEWS-nolinks.html certain substrings contain sequences of word
names. This word parses one such sequence with PARSE-NAME, checks
if the found word name has an entry in LINKS, and if so, outputs the
word with an HTML link, otherwise without. The link gets the URL
prefix defined earlier.
: parse-news ( -- )
case
"<code>" string-parse dup 0= ?of 2drop endof
type "<code>" type
"</code>" string-parse dup 0= ?of 2drop endof
`parse-code execute-parsing "</code>" type
next-case ;
The substrings containing the word names are surrounded by
<code>...</code>, so this word searches (parses) for <code>, then
for </code>, and the result of the latter parse (the string between)
is passed to PARSE-CODE with `PARSE-CODE EXECUTE-PARSING. This word
also has to output the result of the first PARSE-STRING and the
strings parsed for.
news `parse-news execute-parsing
Let PARSE-NEWS work on NEWS.
One shortcoming of the current version of this program is that it
produces links for some words in the prose of the older NEWS. This
could be mostly fixed by only putting links on capitalized words (this
is complicted by cases where the HTML encoding of, e.g. ">" as ">"
contains lower-case characters). For now this is low on the ToDo
list.
This program uses the following words that are not standard and have
not been accepted for standardization:
slurp-file: used twice, quite useful.
string-parse (4x) execute-parsing (3x): make it much easier to search
through the files/strings that if we had used SEARCH. What is missing
(also for search), but was not needed here is to search for one of
several strings; Gforth has a regexp package that may be usable for
that.
?of (5x) next-case (2x): Very useful for defining loops with several
exits. VFX also has these two words. In the present cases the
NEXT-CASE is followed by ";", so some will argue that using a BEGIN
AGAIN loop with IF ... EXIT THEN instead of ?OF ... ENDOF would have
worked just as well, and there is something to it. But sometimes I
want to put a debugging tracer (~~) when a word is exited, and having
several EXITs in addition to ";" means I have to insert several
tracers. Using multiple WHILEs would have resulted in
hard-to-understand code for getting the right stack for each WHILE.
nextname: Often very useful for defining words where you do not have
the names in parseable form. However, looking at the code again, I
actually have the name in parseable form, and if we can live with the empty-parse checking that 2CONSTANT does (and that should be ok for
this program), one can simplify PARSE-LINKDATA as follows:
: parse-linkdata ( -- )
case
"<td class=\"printindex-index-entry\"><a href=\""
string-parse nip 0= ?of endof
"\"><code>" string-parse dup 0= ?of 2drop endof
[: links set-current 2constant ;] current-execute
next-case ;
current-execute: wrappers for preserving certain state that is changed
inside are useful and (if written correctly) prevent the problem that
the programmer writes restoration code that does not work if the
wrapped code throws.
third: can also be written as 2 PICK.
RDROP: can also be written as R> DROP.
.\": can also be written as "<string>" TYPE.
Maybe just as notable is what is not used: a string stack or even
Gforths $tring words. I find that the cases where such features help
in my work are few and far between.
- anton
--
M. Anton Ertl
http://www.complang.tuwien.ac.at/anton/home.html
comp.lang.forth FAQs:
http://www.complang.tuwien.ac.at/forth/faq/toc.html
New standard:
https://forth-standard.org/
EuroForth 2026 CFP:
http://www.euroforth.org/ef26/cfp.html
--- Synchronet 3.22a-Linux NewsLink 1.2