Backend Contracts: fast_* and C++ Utilities

EDI’s statistical model-fitting is implemented once, in C++, under R/EDI/src/, and exposed twice: to R via Rcpp (fast_*_cpp exports) and to Python via pybind11 (edi_kernels, same argument names minus the _cpp suffix, and 0-based instead of 1-based indices — see below). This page documents the conventions and guarantees (and, just as importantly, the non-guarantees) that hold across essentially every one of those backend functions, so individual function docs can link here rather than repeating “not validated by this function” verbatim on every page. See vignette("notation-glossary") for symbol meanings and vignette("reproducibility") for RNG/seed conventions specifically.

The validation boundary: R6 wrapper vs. raw backend

Validation happens in the R6 Design/Inference layer, not in the _cpp backend it calls. A public method like InferenceContinOLS$compute_estimate() runs checkmate assertions (gated by should_run_asserts()/toggle_asserts()) on its own arguments before calling into a fast_*_cpp function — dimension checks, type checks, range checks, NA checks. The backend function itself then does essentially none of that: it trusts the shapes and values it receives. This means:

Argument dimensions and storage order

Numeric domains, overflow, and underflow safeguards

Domain assumptions are almost never checked at the backend level (see “Validation boundary” above) but the numeric kernels themselves do guard against the specific overflow/underflow failure modes their own formulas are prone to:

None of these are “input validation” in the sense of rejecting bad input — they are numerical safety nets that keep a function returning a finite, sane value under the extreme arguments its own optimizer or a diverging model fit can legitimately produce, as distinct from the (mostly absent) checking of whether the input made statistical sense in the first place.

Convergence flags and return-object conventions

Nearly every iterative-fitting backend returns some variant of the same small set of fields — edi::ResultMap’s to_rcpp_list()/to_py_dict() converts whatever fields a given backend .set()s into an R list() or Python dict, so the available fields differ by function, but their meaning, where present, is package-wide:

Field Meaning
converged Logical; whether the optimizer’s stopping criterion was met before maxit was reached. FALSE does not necessarily mean the returned coefficients are useless — it means the convergence tolerance was not certified met, so downstream code should decide whether to trust, retry (e.g. robust_survreg’s random-restart loop), or reject the fit.
iterations Integer count of optimizer iterations actually taken.
gradient_norm The norm of the score/gradient at the returned parameter vector — a continuous convergence diagnostic independent of the boolean converged flag; useful for distinguishing “essentially converged, tolerance was just slightly too tight” from “genuinely still moving.”
neg_loglik / neg_ll / loglik The negative log-likelihood at the fit (neg_loglik/neg_ll are aliases for the same quantity — both are populated so callers used to either naming convention find it) and, where included, loglik = -neg_loglik computed only when finite (NA/omitted otherwise) as a convenience so callers don’t need to negate it themselves.
vcov, std_err, z_vals The parameter covariance matrix, per-parameter standard errors (sqrt(diag(vcov))), and Wald z-statistics (coef / std_err) — present only on backends that were called in variance-computing mode (see estimate_only below).
fisher_information / observed_information / information Three names for the same matrix on backends that expose it (e.g. fast_ordinal_regression_with_var_cpp, fast_cpoisson_combined_with_var, fast_hurdle_negbin_with_var) — populated once, aliased under all three names, since different call sites in the codebase historically settled on different names for the identical quantity.
estimate_only (an argument, not a return field) When TRUE, skips computing the Hessian/Fisher-information-derived quantities above entirely (not just omitting them from the return value) — a real performance path, used e.g. inside bootstrap/randomization inner loops that only need a point estimate per replicate, not a full covariance matrix.

Shared cores and wrapper-to-backend equivalence

Two structural patterns guarantee that “the R version” and “the Python version” of a given kernel are not just similar but the exact same compiled logic:

NA/NaN handling

There is no single package-wide rule for what a backend does when it receives NA/NaN in a numeric argument it does not explicitly document handling for — behavior ranges from “propagates cleanly to a non-finite result field” to “undefined/reads garbage,” and is a direct consequence of the “domains are not checked” rule above. Where a function does have documented NA/NaN handling, it is because that function’s job requires it (e.g. sample_mode_cpp() treats NA — and, for doubles, NaN as a category distinct from NA — as a first-class value that can itself be “the mode,” because computing a mode over data that may contain missingness is exactly the use case that function exists for). Do not assume a fast_*_cpp fitting backend will produce a clean NA in its output merely because its input contained one; check the specific function’s own documentation.