1. CpG_aggregation
1.1. Overview
CpG_aggregation aggregates CpG methylation values over user-defined
genomic regions such as promoters, CpG islands, exons, or other BED intervals.
Two input data types are supported:
count– methylated and total read counts, for example3,10beta– Beta-values, for example0.30
The CpG methylation value must be stored in column 4 of the CpG BED file.
1.2. Input Files
1.2.1. CpG file
The CpG input must be BED3+ with the methylation value in column 4. Compressed input is supported.
Count-data example:
chr1 100017748 100017749 3,10
chr1 100017769 100017770 0,10
chr1 100017853 100017854 16,21
For count input, values must satisfy:
methylated reads >= 0
total reads >= 0
methylated reads <= total reads
Beta-value example:
chr1 100017748 100017749 0.30
chr1 100017769 100017770 0.05
chr1 100017853 100017854 0.76
For beta input, finite values must fall within [0, 1].
1.2.2. Region file
The region file must be BED3+ and may contain any number of additional columns. All original region columns are preserved in the output.
Example:
chr1 567292 568293
chr1 713567 714568
chr1 762401 763402
1.3. Count-mode Outlier Filtering
In count mode, CpG-level outliers are filtered by default.
For each region, the original regional methylation proportion is calculated from all overlapping CpGs with total read count greater than zero:
where \(m_i\) is the methylated-read count and \(n_i\) is the total read count for CpG \(i\).
Each CpG is then tested against this regional methylation probability using a two-sided exact binomial test. A Bonferroni-adjusted cutoff is used:
CpGs with \(p < p_{\mathrm{cutoff}}\) are removed before the filtered regional methylation value is calculated.
Use --no_outlier_filter to disable this filtering.
Outlier filtering is applied only in count mode.
1.4. Usage
Count data:
CpG_aggregation \
-i test_03_RRBS.bed.gz \
-b hg19.RefSeq.union.1Kpromoter.bed.gz \
-t count \
-o out
Beta-values:
CpG_aggregation \
-i methylation.beta.bed.gz \
-b regions.bed.gz \
-t beta \
-o out
Useful options include:
-t,--type {count,beta}– methylation data type in column 4; required-a,--alpha– family-wise error rate for count-mode outlier filtering (default: 0.05)--no_outlier_filter– disable count-mode outlier filtering--min_cpg– minimum number of valid overlapping CpGs required for a region (default: 1)--ddof {0,1}– degrees-of-freedom adjustment used for the standard deviation in Beta-value mode (default: 0)--header– write a header row--na_rep– text used when summary values are unavailable (default:NA)-o,--out_prefix,--output– output prefix
Display all options with:
CpG_aggregation -h
1.5. Output
For output prefix out, the command writes:
out.aggregation.tsv
All columns from the region file are retained, and summary columns are appended.
1.5.1. Count-mode Output
The following columns are appended:
Column |
Description |
|---|---|
|
Number of CpGs retained after outlier filtering. |
|
Total methylated reads after filtering. |
|
Total reads after filtering. |
|
Aggregated methylation proportion after filtering. |
|
Number of valid overlapping CpGs before filtering. |
|
Total methylated reads before filtering. |
|
Total reads before filtering. |
|
Aggregated methylation proportion before filtering. |
CpGs with total read count equal to zero are not included in count-mode aggregation.
1.5.2. Beta-mode Output
The following columns are appended:
Column |
Description |
|---|---|
|
Number of overlapping Beta-values used. |
|
Mean Beta-value. |
|
Median Beta-value. |
|
Minimum Beta-value. |
|
Maximum Beta-value. |
|
Standard deviation of Beta-values. |
1.6. Header and Missing Values
By default, no header is written. With --header, the original region
columns are named:
chrom, start, end, column_4, column_5, …
followed by the aggregation columns.
If a region has fewer than --min_cpg valid overlapping CpGs, the appended
summary fields are written using --na_rep.
Invalid region records or records with an inconsistent number of columns are preserved in the output, with missing summary values appended.