Simulation
CausalStructures.generate_graph — Function
generate_graph([rng], n; m=nothing, p=nothing, class=DAG, latents=0) -> CausalGraphGenerate a random graph on n observed nodes named V1, …, Vn.
Exactly one of m (exact edge count) or p (edge probability) must be given. Edges are sampled uniformly at random over all possible pairs using Erdős–Rényi and oriented by a random topological ordering, guaranteeing acyclicity. This samples uniformly over graphs of a given size/density, but not uniformly over the space of DAGs (some DAGs correspond to more topological orderings than others). For an exact uniform draw from all labelled DAGs on n nodes regardless of edge count, use uniform_dag instead.
class may be DAG (default), CPDAG, ADMG, MAG, or PAG. For CPDAG the sampled DAG is converted via dag_to_cpdag. For ADMG/MAG/PAG, latents additional nodes L1, …, Lk are added to the underlying random DAG and then marginalized out via latent_project; m/p are applied to the full n + latents node set before projection. latents must be 0 (the default) for class = DAG or class = CPDAG.
For class = MAG or class = PAG, each candidate is checked and, if invalid, resampled (up to an internal retry limit) until a valid MAG is found; an error is raised if none is found, which can happen for very dense graphs with many latents. For class = PAG, the resulting MAG is then converted to its equivalence class via mag_to_pag.
rng defaults to Random.default_rng() when omitted; pass an explicit AbstractRNG (e.g. Random.Xoshiro(seed)) for reproducibility.
Examples
dag = generate_graph(7; m = 5)
cpdag = generate_graph(7; m = 5, class = CPDAG)
admg = generate_graph(15; p = 0.25, class = ADMG, latents = 3)
mag = generate_graph(15; p = 0.25, class = MAG, latents = 3)
pag = generate_graph(15; p = 0.25, class = PAG, latents = 3)CausalStructures.uniform_dag — Function
uniform_dag([rng], n::Integer; counts = nothing) -> DAGGenerate a random DAG on n observed nodes named V1, …, Vn uniformly at random from the space of all labelled DAGs on n nodes.
Unlike generate_graph, which samples edges independently (uniform over graphs of a given size/density but not uniform over the space of DAGs), this uses the recursive-enumeration algorithm of Kuipers and Moffa (2015).
rng defaults to Random.default_rng() when omitted; pass an explicit AbstractRNG (e.g. Random.Xoshiro(seed)) for reproducibility.
Cost grows quickly with n (well past O(n^3), since the DP table's BigInt entries also grow in digit-width), unlike generate_graph's roughly edge-linear cost; pass a table from uniform_dag_counts via counts to avoid recomputing it when drawing many DAGs at the same n (or below).
Examples
dag = uniform_dag(6)References
CausalStructures.uniform_dag_counts — Function
uniform_dag_counts(n::Integer) -> Vector{Vector{BigInt}}Precompute the DP table uniform_dag needs to sample DAGs on n nodes. table[m][k] is the number of labelled DAGs on m nodes with exactly k outpoints, for every 1 <= m <= n; entry m only depends on entries < m, so a table computed for some n is also valid for any n' <= n.
Pass the result as uniform_dag(rng, n; counts = table) to avoid recomputing it on every draw when sampling many DAGs at the same (or a smaller) n.
Examples
julia> table = uniform_dag_counts(20);
julia> using Random
julia> dags = [uniform_dag(Xoshiro(i), 20; counts = table) for i = 1:5];References
CausalStructures.simulate_data — Function
simulate_data([rng], cg::DAG; samples, standardize=true,
coef_range=(-1.0, 1.0), error_sd=1.0)
-> Dict{Symbol, Vector{Float64}}Simulate observational data from a linear Gaussian structural causal model over cg. Each node is a linear function of its parents plus independent Gaussian noise with standard deviation error_sd. Edge coefficients are drawn uniformly from coef_range. If standardize = true (default), each variable is standardized to zero mean and unit variance.
rng defaults to Random.default_rng() when omitted; pass an explicit AbstractRNG (e.g. Random.Xoshiro(seed)) for reproducibility.
Returns a Dict mapping each node name to a length-samples vector.
Examples
julia> dag = DAG("A --> B --> C");
julia> data = simulate_data(dag; samples = 100);
julia> haskey(data, :A)
true
julia> length(data[:B])
100