Chromnitron Downstream Analysis Summary cGAS KO vs Control | 16 & 18 CAPs | hg38
Study design: cGAS_KO_neg_Tau vs CLT_neg_Tau (cGASKO_vs_con), 4 replicates per condition. (The companion
TAU_vs_con comparison exists in this pipeline but is excluded from this report by request.)
Two CAP sets compared throughout this report: 16 CAPs (CREB1, EP300, NEUROD1, FUS, CEBPA, CEBPB, CEBPD, CTCF,
CTCFL, SMARCA4, SMARCB1, IRF3, TFE3, MITF, TP53, CGAS) and 18 CAPs (+ TFEB, SMARCA5).
Pipeline: S1 loci selection (CAP-independent, shared by both CAP sets) → S3 inference merge → S4 data matrices → S5 postprocessing → S6 NMF
signatures (k=4) → pathway enrichment (Enrichr + clusterProfiler) → motif analysis.
Each section below has its own CAP set (and, where relevant, background/direction/CAP) selector, so different
sections can be viewing different selections at once.
S1 — Gene Locus Selection
Loci selection is driven by ATAC-seq differential accessibility, independent of which CAPs are later scored — this
step is identical for the 16-CAP and 18-CAP analyses (not rerun per CAP set).
Condition
Direction
Genes
Merged Loci
cGASKO_vs_con
Upregulated
5,468
4,180
cGASKO_vs_con
Downregulated
3,616
cGASKO_vs_con
Conserved
375
S5 — CAP Differential Variability
Applies to this section only.
CAP Track Plots — Differential Binding with Inflammation Gene Annotations
How to read: Each horizontal track = one CAP. X-axis = gene loci (ordered as in the data matrix).
Blue fill = CAP more active in experimental vs control (positive log2FC); red fill = less active (negative log2FC).
Dashed vertical lines mark loci of inflammation-related genes, color-coded by category.
cGASKO vs Control — UpregulatedDifferential CAP binding across gene loci (16cap). Dashed lines = inflammation-related genes.
cGASKO vs Control — DownregulatedDifferential CAP binding across gene loci (16cap). Dashed lines = inflammation-related genes.
cGASKO vs Control — UpregulatedDifferential CAP binding across gene loci (18cap). Dashed lines = inflammation-related genes.
cGASKO vs Control — DownregulatedDifferential CAP binding across gene loci (18cap). Dashed lines = inflammation-related genes.
CAP Variability Bar Plots (per-condition, S5 native output)
cGASKO vs Control — upregulated — CAP variability ranking (16cap)
cGASKO vs Control — downregulated — CAP variability ranking (16cap)
cGASKO vs Control — upregulated — CAP variability ranking (18cap)
cGASKO vs Control — downregulated — CAP variability ranking (18cap)
Why this section exists: in the S5 variability ranking above, CGAS often ranks low despite being the literal
knockout target — its predicted binding at the ATAC-selected candidate genes is sparse in both conditions.
Scanning the whole genome (all ~2.87M x 1kb windows, chr1-22) tells a different story. This scan's result is
numerically identical between CAP sets (CGAS predictions come from the same underlying per-CAP inference
regardless of which CAP subset is analyzed downstream, verified pixel-identical below) — the plot still switches
with this section's CAP set dropdown above so its title always matches the selected CAP set.
Left: genome-wide density of con vs cGASKO predicted CGAS binding per 1kb window. Middle: distribution of the signed change (cGASKO−con). Right: per-chromosome mean, con vs cGASKO. (16cap — data is numerically identical to the other CAP set; verified pixel-identical apart from the title text.)
Left: genome-wide density of con vs cGASKO predicted CGAS binding per 1kb window. Middle: distribution of the signed change (cGASKO−con). Right: per-chromosome mean, con vs cGASKO. (18cap — data is numerically identical to the other CAP set; verified pixel-identical apart from the title text.)
Result: a real, consistent, genome-wide decrease. CGAS shows nonzero predicted binding at 93.3% of all
genome-wide 1kb windows in both conditions, with a clear net decrease upon knockout: mean 0.196 (con) vs 0.152
(cGASKO) — 82.2% of windows decreasing vs only 9.4% increasing. This holds in 22 of 22
chromosomes individually.
Bottom line: the S5 variability plot's low CGAS rank is an artifact of restricting to ATAC-selected loci,
not evidence the model fails to detect the knockout.
S6 — NMF Chromatin Signatures
Applies to this section only.
Method: NMF signature decomposition is run with n_components=4 (k=4), read from each run's own
per-run config file, consistently across both CAP sets and both directions.
NMF CAP Groupings (real cap_group.csv membership, this run)
What this shows: For each CAP independently (cGASKO_vs_con), the top-150 genes ranked by |differential|
(cGASKO − control CAP-specific binding score at that locus) were submitted to Enrichr (GO Biological
Process, KEGG, Reactome) via the enrichR R package, with real whole-genome / matched-pool background
support. Pick a CAP below (or leave it on "All CAPs" for the heatmap view) — the plot and table both update to
match, always plot-first-then-table.
Per-CAP pathway hits, real Enrichr results, ranked by FDR
Applies to this section only.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 16cap, background=whole_genome, direction=combined.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 16cap, background=whole_genome, direction=upregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 16cap, background=whole_genome, direction=downregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 16cap, background=matched_pool, direction=combined.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 16cap, background=matched_pool, direction=upregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 16cap, background=matched_pool, direction=downregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 18cap, background=whole_genome, direction=combined.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 18cap, background=whole_genome, direction=upregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 18cap, background=whole_genome, direction=downregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 18cap, background=matched_pool, direction=combined.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 18cap, background=matched_pool, direction=upregulated.
Per-CAP pathway enrichment heatmap, top 25 terms by p-value — 18cap, background=matched_pool, direction=downregulated.
Applies to this section, and the NMF-derived part of Inflammation-Related Pathways below.
Two ways of forming CAP/gene groups are available via the Grouping method dropdown above.
By condition (default, matches the original pipeline): NMF run separately per celltype (con, cGASKO) on
each celltype's own absolute per-CAP binding matrix (S5-filtered and percentile-sparsified — top/bottom
10% per row and column kept, rest zeroed — not untouched raw signal), 8 groups per CAP set (2 directions x 4 groups), using
only the cGASKO celltype's fit. By differential (new): a single joint SVD decomposition of the
cGASKO−con differential matrix (both ATAC directions pooled — sign is intrinsic to each gene, no
separate up/down split), 4 groups per CAP set, each gene/CAP jointly assigned like NMF's own W/H factorization but
signed-value-compatible. Two independent enrichment methods shown for both, both genuinely background-aware:
Enrichr (bar plots, top 10/group, real background= support) and clusterProfiler ORA (table below,
native universe= support).
NMF-group Enrichr enrichment, top 10 hits/group — 16cap, background=whole_genome.
NMF-group Enrichr enrichment, top 10 hits/group — 16cap, background=matched_pool.
NMF-group Enrichr enrichment, top 10 hits/group — 18cap, background=whole_genome.
NMF-group Enrichr enrichment, top 10 hits/group — 18cap, background=matched_pool.
Differential-clustering (SVD) group enrichment, combined genes, top 10/group — 16cap, background=whole_genome (Enrichr).
Differential-clustering (SVD) group enrichment, positive genes, top 10/group — 16cap, background=whole_genome (Enrichr).
Differential-clustering (SVD) group enrichment, negative genes, top 10/group — 16cap, background=whole_genome (Enrichr).
Differential-clustering (SVD) group enrichment, combined genes, top 10/group — 16cap, background=matched_pool (Enrichr).
Differential-clustering (SVD) group enrichment, positive genes, top 10/group — 16cap, background=matched_pool (Enrichr).
Differential-clustering (SVD) group enrichment, negative genes, top 10/group — 16cap, background=matched_pool (Enrichr).
Differential-clustering (SVD) group enrichment, combined genes, top 10/group — 18cap, background=whole_genome (Enrichr).
Differential-clustering (SVD) group enrichment, positive genes, top 10/group — 18cap, background=whole_genome (Enrichr).
Differential-clustering (SVD) group enrichment, negative genes, top 10/group — 18cap, background=whole_genome (Enrichr).
Differential-clustering (SVD) group enrichment, combined genes, top 10/group — 18cap, background=matched_pool (Enrichr).
Differential-clustering (SVD) group enrichment, positive genes, top 10/group — 18cap, background=matched_pool (Enrichr).
Differential-clustering (SVD) group enrichment, negative genes, top 10/group — 18cap, background=matched_pool (Enrichr).
The Summary at the bottom of this report discusses a TOR/TORC1/mTORC1/lysosome/autophagy signal in the
TFE3/MITF-containing NMF group (group identified automatically — whichever NMF group contains TFE3, per CAP
set/direction, not hand-picked). The three keyword categories are not equally strong:
mTORC1 / TORC1 / TOR signaling — robust, significant in both Enrichr and clusterProfiler, both CAP sets
(e.g. Enrichr "Amino Acids Regulate mTORC1" p=5.6×10-7, FDR=2.2×10-4 (16cap) /
p=2.1×10-7, FDR=7.6×10-5 (18cap);
clusterProfiler "mTOR signaling pathway" p=3.4×10-5, FDR=0.0073 (16cap) /
p=1.1×10-4, FDR=0.0133 (18cap)).
Lysosome — real but weaker and tool-dependent: significant in Enrichr, not in clusterProfiler for the
same GO term (raw p-values and FDR shown below).
Autophagy — nominal at best in either tool (best hit: Reactome "Macroautophagy" p=0.0036, FDR=0.207
(16cap); GO_BP "Positive Regulation Of Autophagy" p=0.0016, FDR=0.091 (18cap), never <0.05 FDR);
for 18cap, KEGG's "Autophagy" and "Lysosome" terms have zero gene overlap with this group and aren't returned by
Enrichr at all.
Mechanistic reason these overlap: LAMTOR1, LAMTOR3, FLCN, NPRL3, and DEPDC5 — the Ragulator/GATOR complex
that physically anchors and regulates mTORC1 on the lysosomal membrane — drive most of both the mTORC1 and
lysosome hits. This is largely one signal (Ragulator-complex genes) read out two ways, not three independent
pathways.
Lysosome evidence (whole-genome background, GO:0032418 "Lysosome Localization", group cGASKO_up_G3):
Tool
CAP set
p-value
FDR
Overlap genes
Enrichr
16cap
8.25×10-5
0.0177 (sig.)
FLCN;VPS33A;PIK3CD;LAMTOR1
Enrichr
18cap
4.28×10-5
0.0081 (sig.)
FLCN;VPS33A;PLEKHM2;LAMTOR1
clusterProfiler
16cap
1.10×10-3
0.207 (n.s.)
LAMTOR1;VPS33A;FLCN;PIK3CD
clusterProfiler
18cap
6.23×10-4
0.119 (n.s.)
LAMTOR1;VPS33A;FLCN;PLEKHM2
Both tools test essentially the same gene overlap and get similar raw p-values, but disagree on FDR
significance because they correct across different term universes (Enrichr: per-library correction; clusterProfiler: per-analysis correction).
For comparison — CTCF/CTCFL group (G1): by raw (uncorrected) p-value, this group also shows a
lysosome-related hit — "Regulation Of Lysosome Size" (GO:0062196), p=7.8×10-4 (whole-genome
background) — nominally significant (p<0.001, well below the conventional 0.05 threshold). This is a genuine
positive signal at the raw-p level, not noise dressed up.
Three things distinguish it from the G3/TFE3-MITF signal above, though: (1) it is driven by only 2 overlapping
genes (SNAPIN, BORCS6) out of a 7-gene term, so the p-value is sensitive to small changes in overlap; (2) it does
not survive FDR correction for the number of pathways tested in the same analysis (FDR=0.24 whole-genome /
0.50 matched-pool background); and (3) it is not corroborated by clusterProfiler for the same term, nor by the
CLEAR motif scan (con-vs-cGASKO Fisher test not significant for this group in any promoter window, and the small
effect that does exist runs in the opposite direction from what a CTCF-driven story would predict).
Read this as: a real, nominally-significant, hypothesis-generating association for G1/CTCF-CTCFL — not a
confirmed, FDR-significant finding on the level of G3/TFE3-MITF.
Full keyword-matched hit tables below (all lysosome/autophagy/mTOR/TORC1/vacuole terms, real p-values and
FDR) — both genuinely respond to this section's own background dropdown above: Enrichr results use the
enrichR R package's real background-restriction support (v3.4+, background= parameter,
routed through Enrichr's Speedrichr backend), same as every other Enrichr table in this report.
Differential-Clustering (SVD) Group Signal — TFE3/MITF Group, Detail
Same idea as the condition-based callout above, but for the differential (cGASKO−con) joint-SVD
grouping: whichever of the 4 groups contains TFE3 is identified automatically (not hand-picked), and every hit
below matches lysosome/autophagy/mTOR/TORC1/vacuole keywords, real p-values and FDR. Since this grouping method
pools both ATAC directions and instead splits by loading sign within the group, results are shown per
gene subset (combined / positive-loading / negative-loading) using the Gene subset dropdown above — both
tools genuinely respond to this section's background dropdown.
clusterProfiler ORA — group pathway hits (background-aware, both grouping methods)
Inflammation-Related Pathways
Applies to this section only.
How these were identified: automated keyword filter over the real per-capset NMF-group Enrichr results and
per-CAP Enrichr results (immune, cytokine, inflam*, leukocyte, myeloid, phagocyt*, TGF-, JAK-STAT, T cell,
interferon, chemokine, interleukin, NF-kB, toll-like, complement, innate, microglia, macrophage) — no hand-picking.
Both tables below are background-aware (enrichR real background support via background=, same data
source as the Per-CAP Pathway Enrichment and Lysosome/Autophagy/mTORC1 sections above) and respond to this
section's own CAP set / background dropdown above. The NMF-group table also follows the Grouping
method / Gene subset toggle from the NMF-Group Pathway Enrichment section above.
Two targeted follow-up analyses using real hg38 sequence: (1) does the literal CLEAR element
(5'-GTCACGTGAC-3') appear at the TFE3/MITF and CTCF/CTCFL NMF groups' target promoters, and does that
change between cGAS-intact and cGAS-KO cells; (2) is there a discoverable de novo sequence motif at genome-wide
CGAS-bound regions (STREME + TOMTOM vs JASPAR2022).
Method: for the TFE3/MITF and CTCF/CTCFL NMF groups (dynamically resolved per celltype/direction from each
NMF run's cap_group.csv, since group ids are not stable across independent NMF fits), extracted
each target gene's strand-aware promoter sequence around its TSS and scanned for an exact match to the
palindromic CLEAR element GTCACGTGAC. Three promoter windows are compared via the dropdown above: the original
TSS−500/+500bp, plus two extended windows, TSS−2000/+500bp and TSS−5000/+500bp (TSS recovered
exactly as the midpoint of the original ±500bp window — verified to reproduce the original scan
exactly at the 500/500 setting).
CLEAR Motif Enrichment vs. GC-Matched Genome Background
Why this is a separate test from the con-vs-cGASKO comparison above: the CLEAR element is a short (10bp),
GC-rich, exact-match motif. Naive uniform-random expectation for an exact 10bp string is astronomically low
(~0.1% chance per 1000bp window), but real gene promoters are GC/CpG-enriched relative to bulk genome, so a
random promoter has a much higher baseline hit rate than that naive number suggests. The con-vs-cGASKO test above
only asks "does the rate differ between the two conditions" — it says nothing about whether the motif is
enriched at all. This test asks that second question directly, HOMER-style: is each group's hit rate
higher than a GC-content-matched background of real genome-wide gene promoters (58,567 GENCODE v45 genes, same
window size, 3 background promoters sampled per target promoter from the matching GC bin), via Fisher's exact
test (equivalent to a hypergeometric enrichment test for a single 2×2 table).
Fingerprint of enrichment across every group/direction/celltype/window combination (both CAP sets) —
color = signed -log10(p) (red = enriched vs. background, blue = depleted; white = not significant):
CLEAR motif enrichment vs. GC-matched genome-wide background — 16cap.
CLEAR motif enrichment vs. GC-matched genome-wide background — 18cap.
Result: TFE3/MITF-upregulated promoters are strongly, unambiguously enriched for the CLEAR motif relative
to GC-matched background in every window and both CAP sets (e.g. 18cap con-upregulated: 13.97% vs. 0.56%
background, p=3.7×10-13) — confirming the motif is genuinely concentrated in this group's
promoters, not just a background-rate coincidence. This holds in both con and cGASKO cells, which is exactly why
the con-vs-cGASKO comparison above isn't significant: cGAS knockout doesn't change an already-strong,
near-saturating signal. CTCF/CTCFL and the downregulated direction show no enrichment above background anywhere
(most p>0.2, several with target rate below background) — a stronger, cleaner negative than "doesn't
survive FDR correction" alone.
HOMER Known-Motif Validation
Method: independent cross-validation of the GC-matched-background result above using HOMER's own
known-motif enrichment engine (findMotifsGenome.pl, HOMER 4.10), run on the cluster. Foreground BEDs
are each NMF group's promoter set (TFE3/MITF or CTCF/CTCFL, upregulated or downregulated). Background is a
genome-wide GENCODE v45 promoter BED at the matching window size. HOMER was run with -size given,
a custom known-motif file for exact CLEAR (GTCACGTGAC), -mknown, -nomotif,
and -nlen 0. The custom motif log-odds threshold is 13.0, which is above the
approximate 2-mismatch score and below the perfect-match score, so the HOMER known-motif call is strict for the
exact CLEAR site. Full scope: both directions, all three promoter windows (TSS−500/+500, −2000/+500,
−5000/+500), both CAP sets, both TF groups, both conditions — 48 combinations total.
Result: HOMER agrees with the GC-matched-background Fisher test on every one of the 48 combinations where
both methods produce a P-value (29/48; HOMER reports no P-value for the other 19 — see table footnote).
TFE3/MITF-upregulated is the only significant category in both methods, in all 12 window × CAP set
× condition combinations (HOMER P=1×10-7 to 1×10-22; target hit rates
4.45–19.53% vs. 0.48–0.65% background). CTCF/CTCFL (both directions) and TFE3/MITF-downregulated are
non-significant everywhere in both methods (HOMER P=1.0 throughout). This is full agreement, 12/12 significant
calls and 17/17 non-significant calls among the 29 directly comparable rows — independently confirming the
GC-matched-background conclusion above with a completely different statistical engine.
NMF group
Direction
Window
CAP set
Condition
Target promoters
Target %
Background %
P-value
FDR
CTCF/CTCFL G1
upregulated
tss-500+500
16cap
cGASKO
234
1.28%
0.61%
1e0
0.1725
CTCF/CTCFL G3
downregulated
tss-500+500
16cap
cGASKO
122
0.00%*
—
—
—
CTCF/CTCFL G1
upregulated
tss-500+500
16cap
con
174
0.00%*
—
—
—
CTCF/CTCFL G2
downregulated
tss-500+500
16cap
con
387
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-500+500
16cap
cGASKO
247
4.45%
0.48%
1e-7
0.0000
TFE3/MITF G1
downregulated
tss-500+500
16cap
cGASKO
311
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-500+500
16cap
con
257
6.61%
0.48%
1e-13
0.0000
TFE3/MITF G1
downregulated
tss-500+500
16cap
con
164
0.61%
0.29%
1e0
0.3804
CTCF/CTCFL G1
upregulated
tss-500+500
18cap
cGASKO
220
1.36%
0.61%
1e0
0.1537
CTCF/CTCFL G3
downregulated
tss-500+500
18cap
cGASKO
124
0.00%*
—
—
—
CTCF/CTCFL G1
upregulated
tss-500+500
18cap
con
209
0.00%*
—
—
—
CTCF/CTCFL G2
downregulated
tss-500+500
18cap
con
433
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-500+500
18cap
cGASKO
215
6.98%
0.56%
1e-11
0.0000
TFE3/MITF G1
downregulated
tss-500+500
18cap
cGASKO
345
0.29%
0.43%
1e0
0.7702
TFE3/MITF G3
upregulated
tss-500+500
18cap
con
158
10.76%
0.41%
1e-18
0.0000
TFE3/MITF G1
downregulated
tss-500+500
18cap
con
123
0.81%
0.27%
1e0
0.2806
CTCF/CTCFL G1
upregulated
tss-2000+500
16cap
cGASKO
234
1.28%
0.61%
1e0
0.1706
CTCF/CTCFL G3
downregulated
tss-2000+500
16cap
cGASKO
122
0.00%*
—
—
—
CTCF/CTCFL G1
upregulated
tss-2000+500
16cap
con
174
0.00%*
—
—
—
CTCF/CTCFL G2
downregulated
tss-2000+500
16cap
con
387
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-2000+500
16cap
cGASKO
247
6.07%
0.52%
1e-11
0.0000
TFE3/MITF G1
downregulated
tss-2000+500
16cap
cGASKO
311
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-2000+500
16cap
con
257
8.17%
0.51%
1e-18
0.0000
TFE3/MITF G1
downregulated
tss-2000+500
16cap
con
164
0.61%
0.35%
1e0
0.4405
CTCF/CTCFL G1
upregulated
tss-2000+500
18cap
cGASKO
220
1.36%
0.62%
1e0
0.1574
CTCF/CTCFL G3
downregulated
tss-2000+500
18cap
cGASKO
124
0.00%*
—
—
—
CTCF/CTCFL G1
upregulated
tss-2000+500
18cap
con
209
0.00%*
—
—
—
CTCF/CTCFL G2
downregulated
tss-2000+500
18cap
con
433
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-2000+500
18cap
cGASKO
215
9.30%
0.57%
1e-17
0.0000
TFE3/MITF G1
downregulated
tss-2000+500
18cap
cGASKO
345
0.29%
0.43%
1e0
0.7725
TFE3/MITF G3
upregulated
tss-2000+500
18cap
con
158
13.29%
0.49%
1e-22
0.0000
TFE3/MITF G1
downregulated
tss-2000+500
18cap
con
123
0.81%
0.34%
1e0
0.3431
CTCF/CTCFL G1
upregulated
tss-5000+500
16cap
cGASKO
234
1.28%
0.74%
1e0
0.2532
CTCF/CTCFL G3
downregulated
tss-5000+500
16cap
cGASKO
122
0.00%*
—
—
—
CTCF/CTCFL G1
upregulated
tss-5000+500
16cap
con
174
0.00%*
—
—
—
CTCF/CTCFL G2
downregulated
tss-5000+500
16cap
con
387
0.26%
0.69%
1e0
0.9300
TFE3/MITF G3
upregulated
tss-5000+500
16cap
cGASKO
247
6.07%
0.63%
1e-10
0.0000
TFE3/MITF G1
downregulated
tss-5000+500
16cap
cGASKO
311
0.00%*
—
—
—
TFE3/MITF G3
upregulated
tss-5000+500
16cap
con
257
8.17%
0.65%
1e-16
0.0000
TFE3/MITF G1
downregulated
tss-5000+500
16cap
con
164
0.61%
0.52%
1e0
0.5723
CTCF/CTCFL G1
upregulated
tss-5000+500
18cap
cGASKO
220
1.36%
0.75%
1e0
0.2277
CTCF/CTCFL G3
downregulated
tss-5000+500
18cap
cGASKO
124
0.00%*
—
—
—
CTCF/CTCFL G1
upregulated
tss-5000+500
18cap
con
209
0.00%*
—
—
—
CTCF/CTCFL G2
downregulated
tss-5000+500
18cap
con
433
0.23%
0.69%
1e0
0.9491
TFE3/MITF G3
upregulated
tss-5000+500
18cap
cGASKO
215
9.30%
0.68%
1e-16
0.0000
TFE3/MITF G1
downregulated
tss-5000+500
18cap
cGASKO
345
0.29%
0.57%
1e0
0.8606
TFE3/MITF G3
upregulated
tss-5000+500
18cap
con
158
13.29%
0.66%
1e-20
0.0000
TFE3/MITF G1
downregulated
tss-5000+500
18cap
con
123
0.81%
0.49%
1e0
0.4504
* 19/48 rows (greyed out): HOMER's internal findKnownMotifs.pl hits a division-by-zero and reports no P-value whenever the target set has exactly 0 promoters with the motif — a known HOMER limitation for null results, not a pipeline issue. Every one of these rows has n_target_hit=0 confirmed directly from the exact-match scan, and all are non-significant (p=0.21–1.0) in the independent GC-matched-background Fisher test above, so this does not affect any conclusion.
De Novo Motif Discovery — Gene-Promoter CGAS Binding (S1 candidate loci)
STREME E-values, top 10 motifs, nonzero- vs zero-CGAS promoters — 16cap. Red = matches a known JASPAR TF.
STREME E-values, top 10 motifs, nonzero- vs zero-CGAS promoters — 18cap. Red = matches a known JASPAR TF.
De Novo Motif Discovery — Genome-Wide CGAS Binding (not ATAC-restricted)
Genome-wide scan (chr1-22, no ATAC pre-filter): real per-basepair merged inference output max-pooled per 1kb window,
2 candidate sets each vs. a random zero-signal background (ENCODE blacklist excluded, ≥5kb spacing): (1) top-200
windows by raw CGAS binding in cGAS-intact (con) cells; (2) top-200 windows by |cGASKO−con| differential CGAS
binding. TOMTOM annotation against the full JASPAR2022 CORE non-redundant vertebrate motif database (1,956 motifs).
STREME E-values, top 10 motifs per analysis — 16cap. Red = matches a known JASPAR TF (labeled); blue = below E=0.05 but unmatched; grey = not significant.
STREME E-values, top 10 motifs per analysis — 18cap. Red = matches a known JASPAR TF (labeled); blue = below E=0.05 but unmatched; grey = not significant.
Follow-Up on the Genome-Wide CGAS-Decrease Regions — Genes & Pathways
Method: for every gene TSS in GENCODE v45 (58,567 loci genome-wide, not the S1 candidate list), looked up the
predicted CGAS binding score at its own TSS-containing 1kb window in both con and cGASKO, computed con−cGASKO.
Ranked all genes by this signed decrease. Rank-based GSEA (GO_BP + KEGG, full ranked list) and threshold ORA
(GO_BP, top-200 by decrease and top-200 by increase, universe = all TSS-scored genes) both run per CAP set (this
analysis is CAP-set-independent in its CGAS score, since CGAS predictions don't depend on which other CAPs are
analyzed — but is repeated per CAP set here for a complete, self-contained artifact).
ORA (clusterProfiler enrichGO), top-200 genes by TSS CGAS decrease / increase — 16cap.
ORA (clusterProfiler enrichGO), top-200 genes by TSS CGAS decrease / increase — 18cap.
GSEA (GO Biological Process, full ranked list) — top 10 by FDR — 16cap
Term
NES
p-value
FDR
Set size
muscle tissue development
1.49
1.03e-08
6.39e-05
404
cardiac septum development
1.73
1.07e-07
1.33e-04
106
cardiac chamber development
1.65
8.72e-08
1.33e-04
166
double-strand break repair
1.53
6.56e-08
1.33e-04
297
regulation of amide metabolic process
1.44
4.46e-08
1.33e-04
458
transmembrane receptor protein serine/threonine kinase signaling pathway
1.47
1.72e-07
1.78e-04
357
epigenetic regulation of gene expression
1.56
2.62e-07
1.81e-04
215
striated muscle tissue development
1.54
2.44e-07
1.81e-04
242
proteasome-mediated ubiquitin-dependent protein catabolic process
1.42
2.08e-07
1.81e-04
444
cell growth
1.42
2.99e-07
1.86e-04
467
GSEA (GO Biological Process, full ranked list) — top 10 by FDR — 18cap
Term
NES
p-value
FDR
Set size
muscle tissue development
1.49
1.03e-08
6.39e-05
404
cardiac septum development
1.73
1.07e-07
1.33e-04
106
cardiac chamber development
1.65
8.72e-08
1.33e-04
166
double-strand break repair
1.53
6.56e-08
1.33e-04
297
regulation of amide metabolic process
1.44
4.46e-08
1.33e-04
458
transmembrane receptor protein serine/threonine kinase signaling pathway
1.47
1.72e-07
1.78e-04
357
epigenetic regulation of gene expression
1.56
2.62e-07
1.81e-04
215
striated muscle tissue development
1.54
2.44e-07
1.81e-04
242
proteasome-mediated ubiquitin-dependent protein catabolic process
1.42
2.08e-07
1.81e-04
444
cell growth
1.42
2.99e-07
1.86e-04
467
Summary — Both CAP Sets
Both CAP sets reproduce the same mTORC1/Ragulator signal — lysosome is real but weaker, autophagy is nominal only
The TFE3/MITF-containing NMF group (cGASKO_up_G3 in both CAP sets) shows a robust TOR/TORC1/mTORC1 signaling
hit in both Enrichr and clusterProfiler (e.g. Enrichr "Amino Acids Regulate mTORC1" FDR=2.2×10-4
(16cap) / 7.6×10-5 (18cap); clusterProfiler "mTOR signaling pathway" FDR=0.0073 (16cap) / 0.0133
(18cap)). Lysosome is real but weaker and tool-dependent: Enrichr's "Lysosome Localization" (GO:0032418) reaches
FDR=0.018 (16cap, p=8.3×10-5) / 0.0081 (18cap, p=4.3×10-5), but the identical term
in clusterProfiler does not clear FDR<0.05 (FDR=0.21 / 0.12). Autophagy terms are nominal at best (best hit
FDR=0.27 (16cap) / 0.091 (18cap), never significant in either tool). These aren't three independent signals —
LAMTOR1, LAMTOR3, FLCN, NPRL3, and DEPDC5 (the Ragulator/GATOR complex that anchors mTORC1 to the lysosomal
membrane) drive most of both the mTORC1 and lysosome hits. See the dedicated "mTORC1 (Ragulator Complex) /
Lysosome / Autophagy Signal — TFE3/MITF Group, Detail" callout in the NMF-Group Pathway Enrichment section
above for the full terms, p-values, FDR values, and genes, per CAP set.
Genome-wide CGAS signal is CAP-set-independent
The genome-wide CGAS binding pattern, gene-level TSS CGAS-decrease ranking, and de novo motif discovery results
are numerically identical between CAP sets (CGAS predictions come from the same underlying per-CAP inference
regardless of which other CAPs are included downstream) — a useful cross-check that both pipeline reruns are
processing the same underlying model output correctly.