From Newsgroup: muc.lists.freebsd.ports
On 9/26/26 16:55, John W. O'Brien wrote:
Hello FreeBSD-Ports,
I switched a few months ago from running my main Poudriere builder on an
old Dell PowerEdge server to running it on an AWS EC2 instance
(c5.12xlarge in us-east-2). The instance was initially deployed from the "FreeBSD 14.4-RELEASE-amd64 UEFI-PREFERRED base ZFS" AMI (ami-075b8663dbd980c35) and has since been updated to 14.5-RELEASE. This instance has a 200G EBS volume of which 180G is a ZFS pool with roughly
80G free and 20G of swap to complement 96G of RAM. The instance does
nothing else besides shipping built packages via Nginx.
With relatively high regularity---most, but not all "bulk" runs---
poudriere will terminate with some variation of the following mix of
errors.
_mktemp: mkstemp failed on /srv/poudriere/data/logs/bulk/145amd64-
default-
general/2026-09-26_22h00m19s/.tmp-.poudriere.snap_loadavg.Ms7EnmmR: No
space left on device
echo: write error on stdout
_mktemp: mkstemp failed on /srv/poudriere/data/logs/bulk/145amd64- default-general/2026-09-26_22h00m19s/.tmp-.poudriere.status.10gWzEcY: No space left on device
err: Recursive error detected: write_atomic unable to create tmpfile
in /srv/poudriere/data/logs/bulk/145amd64-default- general/2026-09-26_22h00m19s
_mktemp: mkstemp failed on /srv/poudriere/data/logs/bulk/145amd64- default-general/2026-09-26_22h00m19s/.tmp-.poudriere.ended.ld5bXm8y: No space left on device
err: Recursive error detected: write_atomic unable to create tmpfile
in /srv/poudriere/data/logs/bulk/145amd64-default- general/2026-09-26_22h00m19s
echo: write error on stdout
This usually leaves one or two builder jails running, and I have gotten
in the habit of rebooting the instance before re-starting the build.
If poudriere cannot shut down gracefully on failure then a bug should
be reported if one doesn't already exist.
A typical build pulls in between 1000 and 5000 ports. I have not
detected a correlation between the crash and a specific port, but I have also not been keeping especially close track. Subjectively, the failure usually occurs in the latter third or latter quarter by port count.
I haven't investigated but I'd assume you could filter through logs
of builds that did not complete/fail and may be able to check on status through the web interface of past builds if setup.
My poudriere.conf is only lightly adapted from stock, and none of the
memory limit parameters are modified from "(default: none)".
I'm not sure what I should be looking at to narrow down the cause of
these failures. Any suggestions?
My TLDR is:
* Take CPU core count and mathematically find its factors (=2 numbers
that multiply to the core count)
example with 32 cores: 1*32, 2*16, 4*8
* /usr/local/etc/poudriere.conf:PARALLEL_JOBS= a factor < the total core
count
* /usr/local/etc/poudriere.conf:ALLOW_MAKE_JOBS=yes
* /usr/local/etc/poudriere.d/make.conf:MAKE_JOBS_NUMBER= the other factor
* Keep adjusting PARALLEL lower and MAKE higher until you are happy with
the way the system is loaded. You can use this combination to set the
system to try to always be over/under-saturated with work if you pick
two #s that are more or less than the core count.
* Some ports do not properly respect MAKE_JOBS_NUMBER. digikam, mongodb,
and deno come to mind but I don't know if any have fixed that. The
result is with MAKE_JOBS_NUMBER=6 I have found that 6*6=36 processes are sometimes running which seems to be an issue of a build system inside a
build system and both receive the same job count.
Poudriere defaults to the corecount factors of
PARALLEL_JOBS=corecount and MAKE_JOBS_NUMBER=1 which can easily create a
worst case RAM use scenario. On my 32GB RAM 4 core + hyperthreaded (=8 "cores") system I normally have PARALLEL=2 and MAKE at 8 but even then I
still need swap to get through builds. I assume official builders have PARALLEL=cores and MAKE=1 and I know I've seen the poudriere page of
official jobs say 100% RAM used and 98% swap used so they likely could
use some tuning themself if that was accurate and still present. I
assume builds could complete faster if we get the package builders to
use a minimal amount of swap.
Some more detailed analysis of options:
inside poudiere.conf:
USE_TMPFS=yes
This default includes writing the port's entire workdir into RAM and as
I understand it its not compressed storage space. If you have over 20GB
of extracted content + compiler output then you are spending a lot of
RAM on a package. I've observed rust at over 30GB but not sure what it currently gets up to. Maybe using ZFS on a memory disk with default compression could reduce memory impact of some work without a noticeable impact to CPU/RAM performance but I haven't benchmarked it.
TMPFS_BLACKLIST
This can be used to force selected ports to lose the TMPFS setting so a candidate like rust which may be too big to handle can have its work
directory go out to disk instead of RAM.
PARALLEL_JOBS
With a default at the number of CPUs, this setting causes multiple ports
to be actively built in parallel. The RAM impact of TMPFS is now the sum
of all ports actively being worked on.
ALLOW_MAKE_JOBS
Disabled by default but when active it permits the build of a single
port to use multiple jobs. I find the use of RAM per make job (compiler processes are a common high RAM draw) is significantly lower than the
TMPFS load of many big ports.
ALLOW_MAKE_JOBS_PACKAGES
A list of ports that have ALLOW_MAKE_JOBS enabled. This then comes into
play with the ports tree's MAKE_JOBS_NUMBER whether you considered it or
not.
MUTUALLY_EXCLUSIVE_BUILD_PACKAGES
A list of ports that will not be built until no other job is currently running. In case you cannot keep something like rust building with
anything else because it alone is too close to max hardware use then add
it here to lower such load.
PACKAGE_FETCH_WHITELIST
If you haven't customized how a port or its dependencies is built that
is used as a dependency to other ports then you can look into a blacklist/whitelist way of permitting some to be fetched from official repositories.
PRIORITY_BOOST
You can set this to a list of ports to focus on first. If you are
working on too many large things all at once then adding them here might change the order so they don't all get worked on at the same time. Using
this is more beneficial as it can help keep more builders busy instead
of a big queue being stuck behind 1 port completing when other unrelated
stuff was already finished first.
Using ccache can sometimes make builds complete faster and whenever its
cache is helping you will also lose the high RAM requirements of running
the compiler for any given task.
If you try to play with settings that take only integers and not decimal numbers, sometimes you can use basic math to scale to what you wanted: MAX_MEMORY="400 / 100"
That would be an overkill way to say 4, but now you can edit the 400 to
other integers to adjust 10ths and 100ths fraction amounts and could
push the math further to reach other values.
I think the TMPFS and MEMORY limit variables cause things exceeding it
to error out rather than be rescheduled/retried so I haven't really
found a good way to use them as I'd rather fix the issue instead of mask
it. Unless the ports tree can tell a scheduler what load to expect for a
port its not likely that poudriere would be making any good scheduling decisions with it anyway.
I've had ideas of how poudriere + the ports tree could try to
optimize job scheduling better but haven't tried to put it into any form
of code; before poudriere replaced tinderbox I had been thinking of ways
to try to write a more efficient build system. There are simple and
exotic ideas where even simple ones would likely help reduce the value
of job boosting and such without much of a chance to schedule anything
worse. Other than having a smarter scheduler that could try to fend off certain bad cases automatically, a better build system is likely to only stress systems more.
Thank you,
John
--
Posted automagically by a mail2news gateway at muc.de e.V.
Please direct questions, flames, donations, etc. to
news-admin@muc.de
--- Synchronet 3.22a-Linux NewsLink 1.2