Options controlling the parallel backend and its messaging, set via
options.
Details
arf.backendParallel backend used when
parallel = TRUE:"foreach"or"mirai". If unset, arf uses"mirai"when mirai daemons are running and"foreach"otherwise.arf.verboseReport the selected backend once per backend configuration per session? Default
TRUE; setFALSEto silence.arf.block_rowsCap on rows materialized per block of conditions in
expct. Default5e6. Lower it to trade speed for memory on large forests with many conditions.
arf.block_rows does not affect results, only peak memory and speed.
For memory-constrained hardware, combine a small daemon count with a lower
arf.block_rows and the batch/stepsize arguments of
lik, forge and expct.
The "foreach" backend uses whatever adapter is registered (e.g.
doParallel, doFuture). The "mirai" backend uses
mirai daemons and shares large read-only inputs (training data,
forest, learned parameters) across workers via mori, so workers do
not each copy them. In our benchmarks (n = 20000, 200 trees, up to 16
workers) run time was similar between the backends, though on a cold daemon
pool "mirai" can be noticeably slower. The memory benefit is largest
for the tree-parallel operations (forde,
adversarial_rf) on large forests with many workers: at 16
workers peak memory was about 30 percent lower for forde and about
half for adversarial_rf. Expect smaller gains on smaller problems,
where the daemon pool's fixed overhead can outweigh the sharing, so prefer
"mirai" at scale and either backend otherwise.
Daemons started via future.mirai (e.g.
plan(future.mirai::mirai_multisession)) are detected like any other
mirai daemons, so futureverse users get the "mirai" backend
automatically. None of these packages are installed with arf: install
mirai and mori for the "mirai" backend,
doParallel or doFuture for "foreach", and
doRNG or future.mirai if you use them below.
Workers and data.table threads multiply. If their product exceeds
the core count, cap threads per worker with
data.table::setDTthreads() and set OMP_WAIT_POLICY=passive
before starting R so idle OpenMP threads sleep instead of spinning.
Reproducibility of stochastic operations (forge, categorical
expct) under parallel execution: set.seed only
governs the calling process, not the workers. With the "mirai"
backend, seed the daemons instead: mirai::daemons(n, seed = 42)
gives reproducible results provided the daemon count, the seed and the
sequence of calls on a fresh daemon pool are kept fixed (changing the
daemon count changes how work is chunked and therefore the random stream
assignment). This also applies to pools started via future.mirai:
future's own seed machinery covers only work dispatched through its
API, which arf's backend bypasses, and plan() does not accept a
seed argument, so seed the pool with mirai::daemons()
directly. For the "foreach" backend, register doRNG on top of
the adapter (doRNG::registerDoRNG(42) after
registerDoParallel() or registerDoFuture()): results are then
reproducible, identical across adapters, and independent of the worker
count, provided stepsize is set explicitly (its default depends on
the worker count). Without doRNG, doFuture flags the
stochastic operations with "UNRELIABLE VALUE" warnings. Sequential
execution (parallel = FALSE) with set.seed is exact as
always.
Examples
if (FALSE) { # \dontrun{
arf <- adversarial_rf(iris)
# foreach backend
doParallel::registerDoParallel(cores = 4)
psi <- forde(arf, iris)
# mirai backend: start daemons, then call as usual
mirai::daemons(4)
psi <- forde(arf, iris)
mirai::daemons(0) # shut down when done
# futureverse: future.mirai daemons are detected automatically
future::plan(future.mirai::mirai_multisession, workers = 4)
psi <- forde(arf, iris)
future::plan("sequential") # shuts the daemons down
# force a backend regardless of what is registered (NULL restores auto-detection)
options(arf.backend = "mirai")
options(arf.backend = NULL)
# silence the backend message
options(arf.verbose = FALSE)
# reproducible parallel sampling with mirai: seeded daemons
# (fixed daemon count, fresh pool)
evi <- data.frame(Species = sample(levels(iris$Species), 100, replace = TRUE))
mirai::daemons(4, seed = 42)
# stepsize = evidence rows per step; 25 = 100 conditions / 4 workers,
# i.e. the default sizing, made explicit
x <- forge(psi, n_synth = 1, evidence = evi, stepsize = 25)
mirai::daemons(0)
# reproducible parallel sampling with foreach: doRNG on top of the
# adapter; set stepsize explicitly (its default depends on worker count)
doParallel::registerDoParallel(cores = 4)
doRNG::registerDoRNG(42)
x <- forge(psi, n_synth = 1, evidence = evi, stepsize = 25)
} # }
