Simulation

CausalStructures.generate_graph — Function
generate_graph([rng], n; m=nothing, p=nothing, class=DAG, latents=0) -> CausalGraph

Generate a random graph on n observed nodes named V1, …, Vn.

Exactly one of m (exact edge count) or p (edge probability) must be given. Edges are sampled uniformly at random over all possible pairs using Erdős–Rényi and oriented by a random topological ordering, guaranteeing acyclicity. This samples uniformly over graphs of a given size/density, but not uniformly over the space of DAGs (some DAGs correspond to more topological orderings than others). For an exact uniform draw from all labelled DAGs on n nodes regardless of edge count, use uniform_dag instead.

class may be DAG (default), CPDAG, ADMG, MAG, or PAG. For CPDAG the sampled DAG is converted via dag_to_cpdag. For ADMG/MAG/PAG, latents additional nodes L1, …, Lk are added to the underlying random DAG and then marginalized out via latent_project; m/p are applied to the full n + latents node set before projection. latents must be 0 (the default) for class = DAG or class = CPDAG.

For class = MAG or class = PAG, each candidate is checked and, if invalid, resampled (up to an internal retry limit) until a valid MAG is found; an error is raised if none is found, which can happen for very dense graphs with many latents. For class = PAG, the resulting MAG is then converted to its equivalence class via mag_to_pag.

rng defaults to Random.default_rng() when omitted; pass an explicit AbstractRNG (e.g. Random.Xoshiro(seed)) for reproducibility.

Examples

dag = generate_graph(7; m = 5)

cpdag = generate_graph(7; m = 5, class = CPDAG)

admg = generate_graph(15; p = 0.25, class = ADMG, latents = 3)

mag = generate_graph(15; p = 0.25, class = MAG, latents = 3)

pag = generate_graph(15; p = 0.25, class = PAG, latents = 3)
source
CausalStructures.uniform_dag — Function
uniform_dag([rng], n::Integer; counts = nothing) -> DAG

Generate a random DAG on n observed nodes named V1, …, Vn uniformly at random from the space of all labelled DAGs on n nodes.

Unlike generate_graph, which samples edges independently (uniform over graphs of a given size/density but not uniform over the space of DAGs), this uses the recursive-enumeration algorithm of Kuipers and Moffa (2015).

rng defaults to Random.default_rng() when omitted; pass an explicit AbstractRNG (e.g. Random.Xoshiro(seed)) for reproducibility.

Cost grows quickly with n (well past O(n^3), since the DP table's BigInt entries also grow in digit-width), unlike generate_graph's roughly edge-linear cost; pass a table from uniform_dag_counts via counts to avoid recomputing it when drawing many DAGs at the same n (or below).

Examples

dag = uniform_dag(6)

References

source
CausalStructures.uniform_dag_counts — Function
uniform_dag_counts(n::Integer) -> Vector{Vector{BigInt}}

Precompute the DP table uniform_dag needs to sample DAGs on n nodes. table[m][k] is the number of labelled DAGs on m nodes with exactly k outpoints, for every 1 <= m <= n; entry m only depends on entries < m, so a table computed for some n is also valid for any n' <= n.

Pass the result as uniform_dag(rng, n; counts = table) to avoid recomputing it on every draw when sampling many DAGs at the same (or a smaller) n.

Examples

julia> table = uniform_dag_counts(20);

julia> using Random

julia> dags = [uniform_dag(Xoshiro(i), 20; counts = table) for i = 1:5];

References

source
CausalStructures.simulate_data — Function
simulate_data([rng], cg::DAG; samples, standardize=true,
              coef_range=(-1.0, 1.0), error_sd=1.0)
    -> Dict{Symbol, Vector{Float64}}

Simulate observational data from a linear Gaussian structural causal model over cg. Each node is a linear function of its parents plus independent Gaussian noise with standard deviation error_sd. Edge coefficients are drawn uniformly from coef_range. If standardize = true (default), each variable is standardized to zero mean and unit variance.

rng defaults to Random.default_rng() when omitted; pass an explicit AbstractRNG (e.g. Random.Xoshiro(seed)) for reproducibility.

Returns a Dict mapping each node name to a length-samples vector.

Examples

julia> dag = DAG("A --> B --> C");

julia> data = simulate_data(dag; samples = 100);

julia> haskey(data, :A)
true

julia> length(data[:B])
100
source