LGASieve

Data availability audit | Version 1

Complete whole-exome files in Dropbox and public repositories

Online-folder audit and file-by-file testing of complete human WES BAM/CRAM downloads, excluding regional slices

Audit date
1 August 2026
Dropbox result
0 complete WES alignments
Public result
13,948 verified indexed alignments
Primary use
Research validation and remote regional analysis
Meaning of "verified downloadable." Every counted endpoint received an anonymous HTTP GET. An object was counted as analysis-ready only when both the alignment and BAI/CRAI index passed the source audit. The full 98.885 TB corpus was not downloaded; verification used real byte transfers, range behavior, metadata sizes, file signatures, and repository checksums where supplied.

The Dropbox files are slices; complete WES files are available publicly

Dropbox. The online /Yang Shao/Personal/Codexhome/Bioinformatics folder contains 342 exact .bam files. All were classified as BRCA mini-BAMs, simulated or derived cohort files, a demo, or spike-ins. The largest is a 649.95 MB SpikeForge product. Searches for complete WES BAM/CRAM, whole-exome FASTQ, and SRA objects found no qualifying file. Full-exome .bai files under _bai are indexes only; their corresponding multi-gigabyte BAMs remain in public repositories.

Public repositories. The audit identified 13,948 repository-described complete human WES alignment objects with a live companion index and 895 additional live WES BAM objects without a submitted index. The indexed collection occupies 98.885 TB (89.935 TiB). The downloadable inventories list every URL, format, size, study, sample accession, confidence class, and verification result.

Scientific relevance. Availability is not BRCA deletion/duplication truth. Only two complete exomes in the searched public corpus have independently supported BRCA exon-level events suitable as positive controls: NA18949 and HG01528, both BRCA1 deletions. No confirmed full-exome BRCA duplication-positive control was found.

Bottom line

Start with assay-compatible 1000 Genomes exomes, the two WGS-supported BRCA1 deletion positives, and the SEQC2 matched tumor/normal replicates. Treat the broad ENA catalog as a discovery pool that still requires header, capture-kit, reference-build, depth, and sample-QC screening.

Hypothesis and counting rules

Hypothesis: the online Dropbox folder contains regional analysis artifacts rather than complete exomes, but complete, anonymously downloadable WES alignments can be identified in official repositories and accessed remotely if both the alignment and its index are live.

ClassOperational definitionTreatment
Indexed complete WESRepository-labelled human WXS BAM/CRAM representing a full run or sample, with a live BAI/CRAI.Primary count
Unindexed complete WESRepository-labelled human WXS BAM with a live endpoint but no submitted BAI.Separate count
Raw WES readsPaired FASTQ data requiring alignment before regional analysis.Separate example cohort
Slice or panelBRCA-only, chromosome-only, gene-panel, mini-BAM, synthetic, spike-in, or demo object.Excluded
Controlled dataLow-level reads requiring dbGaP, EGA, UK Biobank, or other approval.Not downloadable now

"Complete" describes object scope, not clinical adequacy. A repository-labelled WXS object can still be low-yield, poorly covered, tumor-derived, or incompatible with LGASieve's reference and capture design.

Online search, official metadata, and live transfer checks

  1. Audit Dropbox online. Searches were run in the Dropbox website from the Bioinformatics folder, not inferred from the partially synchronized local mirror. Terms included .bam, .fastq.gz, .fq.gz, .cram, .sra, exome, whole exome, and WES.
  2. Enumerate official public sources. File-level inventories came from 1000 Genomes/IGSR, ENA Portal API, DepMap/CCLE public AWS data, NIST GIAB, Google Brain GIAB benchmark data, NCBI SEQC2, and Broad GATK AWS test data.
  3. Exclude fragments. Chromosome-split files, unmapped-read supplements, regional slices, panels, VCF-only products, and controlled-access files were excluded.
  4. Verify every reported URL. Anonymous range GETs were issued to each alignment and index. Response size was checked against official metadata, and reserved URL characters were encoded before retrying.
  5. Challenge range behavior. Seven repository/format representatives were queried at the start, middle, and tail with 64 KiB requests. All 21 returned HTTP 206, exact ranges, distinct payload hashes, and valid BAM/CRAM start signatures. Two positive 1000 Genomes BAMs were also accessed successfully as BRCA-region slices.
  6. Reconcile counts conservatively. Mirrors and exact MD5 aliases were grouped where justified. Technical replicates, capture kits, and reference-build alignments remain separate objects. Similar names or sizes alone were never used to merge data.

Main result: verified indexed files by repository

Repository / collectionFormatVerified indexed objectsSample identifiersIndexed volumeInterpretation
1000 Genomes ProjectBAM2,5512,55125.014 TBOfficial QA-passed Phase 3 WES; preferred first cohort
European Nucleotide Archive4,675 BAM + 5,927 CRAM10,6025,999 in indexed set65.658 TBSubmitter WXS metadata; requires header, capture, and coverage QC
DepMap CCLEBAM4764765.621 TBCancer cell lines; not germline-normal truth
Google Brain GIAB benchmarkBAM26691.210 TBMulti-platform and capture-kit technical replicates
NIST GIAB / NCBIBAM2851.003 TBCurated benchmark samples; repeated technologies
NCBI SEQC2BAM242365.355 GB12 tumor and 12 matched-normal multi-center replicates
Broad GATK AWS test dataBAM1113.105 GBSingle official WES test object
TotalBAM + CRAM13,948Not additive98.885 TBObjects, not independent people

The downloadable CSV is authoritative. Repository-scoped identifiers, file objects, logical alignments, and independent people are different denominators.

Confidence and counting denominators

DenominatorObjectsMeaning
Curated or benchmark WES2,869Official release or purpose-built benchmark; still requires assay compatibility checks
Repository-metadata WXS candidates12,009Must pass header, capture-footprint, and coverage QC before use
Unique server objects14,878Distinct normalized alignment URLs after mirror reconciliation
Conservative logical alignments14,842After 36 exact or source-declared aliases; not a count of people
Download-verified alignments14,84313,948 indexed plus 895 unindexed BAMs
Failed or unverified candidates35Excluded from downloadable counts

Dropbox classification

Online BAM classFilesWhy not complete WES
LGASieve BRCA-region mini-BAM150Only selected BRCA regions
LGASieve synthetic/derived cohort BAM117Generated analysis artifact
Other LGASieve mini-BAM20Regional subset
LGASieve demo BAM1Demonstration input
SpikeForge spike-in BAM54Modified/spiked derivative
Total online BAMs3420 qualifying complete WES files

Verified indexed file-size distribution

Alignment sizeFilesKnown volumeInterpretation
<100 MiB2158.505 GBUnusually small; strict content and coverage review required
100 MiB to <1 GiB5,0433.100 TBOften CRAM or low-yield submissions; verify callable exome footprint
1 to <5 GiB1,8195.887 TBCommon for CRAM and smaller BAMs
5 to <10 GiB3,31226.276 TBCommon full-WES range
10 to <20 GiB3,01243.689 TBLargest share of total volume
20 to <50 GiB45914.306 TBDeep, multi-library, or less compressed alignments
≥50 GiB885.619 TBOutliers requiring sample and library review

Size is not sequencing depth. CRAM can be much smaller than an equivalent BAM because it uses reference-based compression.

Downloadable or candidate data outside the indexed total

ClassFilesWhy separateNext step
ENA WXS BAM without submitted BAI895 verifiedDownload works; regional random access is not readyDownload, validate, and create a local BAI
PRJNA388048 BRCA-relevant raw WES120 FASTQs / 60 runsNo public indexed BAM/CRAM; published CNA candidates lack independent confirmationAlign and QC as a research candidate cohort
ENA unindexed BAM failures8Tiny HTML/directory response or metadata-size mismatchExcluded; retry archive later
Indexed ENA candidate failures27Alignment or index did not resolve to the metadata-sized genomic objectExcluded from all verified totals

Two strong deletion positives; no confirmed duplication positive

SampleComplete WESIndependent evidenceWES depthLGASieve result
NA1894911,042,427,115 B BAM + BAIBRCA1 E14-E15 deletion; Phase 3 PASS SV, later 30x ensemble, and DRAGEN Manta support107.282xCorrect depth candidate; conservative policy returned REVIEW because the exome slice lacked independent breakpoint/allelic support
HG0152810,319,084,612 B BAM + BAIBRCA1 E01-E06 deletion; later 30x ensemble, DRAGEN Manta, and DRAGEN CNV support138.939xCorrect deletion signal; conservative policy returned REVIEW
Do not call the rest true negatives. The 150-sample 1000 Genomes subset has no catalogued exon-overlapping BRCA deletion/duplication and is useful for specificity/referral auditing, but public SV callsets are not exhaustive clinical truth. Two deletion positives and zero duplication positives cannot estimate clinical sensitivity.

SEQC2 is the strongest general CNV stress set found

SEQC2 provides 24 indexed WES BAMs: 12 HCC1395 tumor and 12 matched HCC1395BL normal multi-center replicates, using Agilent SureSelect Human All Exon v6 + UTR and GRCh38.d1.vd1. Its high-confidence somatic CNV set integrates six callers, 21 WGS replicates, and CytoScan, BeadChip, and Bionano evidence. It is useful for reproducibility and gain/loss stress testing, but it is not germline BRCA exon-LGA clinical truth.

Analyze regions remotely; do not copy the full corpus to this computer

The indexed URLs support byte-range access. htslib/samtools-based software can request BRCA1/2 intervals and index blocks while leaving the multi-gigabyte alignment in the repository.

  1. First: 1000 Genomes GRCh37. Use the 2,551 official exomes because their build matches the current target model. Retain the 150-sample stratum for specificity/referral work and add NA18949 and HG01528 as deletion controls.
  2. Second: SEQC2. Run the 12 tumor/normal replicate pairs to test cross-center reproducibility and gain/loss stability against published CNV truth.
  3. Third: capture-matched GIAB data. Use multi-kit and multi-platform replicates to quantify capture-kit and alignment-build effects. Do not count GRCh37 and GRCh38 versions as independent people.
  4. Fourth: selected ENA/DepMap cohorts. Inspect headers, target design, read groups, reference, sample type, and depth. Build panels of normals by capture kit and laboratory.
  5. Unindexed BAMs. Use only when phenotype or truth is valuable enough to justify full download and local indexing.

What this audit proves, and what it does not

  • Endpoint access, not full-file integrity. Every counted URL was exercised with a real GET, but the entire 98.885 TB was not transferred and re-checksummed.
  • WXS metadata is not clinical QC. ENA labels are submitter-declared. Small objects are retained and flagged, not asserted to have an adequate callable exome.
  • Counts are time-stamped. Archives change; the scripts and query definitions permit refresh.
  • Controlled cohorts are absent. TCGA/GDC, EGA, UK Biobank, All of Us, dbGaP, and similar low-level human data could not satisfy anonymous-download verification.
  • Availability is not clinical validation. Blinded, assay-matched, orthogonally confirmed deletion and duplication cohorts remain necessary.

Direct answers

  1. Complete WES in the online Dropbox folder? None found. The 342 BAMs are regional, derived/synthetic, demo, or spike-in objects.
  2. Complete exomes elsewhere? Yes: 13,948 indexed public WES alignment objects plus 895 downloadable unindexed BAMs.
  3. Can they be used without downloading whole files? The indexed objects support remote ranges. CRAM also requires the exact reference. Unindexed BAMs do not support efficient regional use.
  4. Can all validate BRCA sensitivity? No. Most lack independent BRCA exon-CNV truth. The two deletion controls and SEQC2 are immediate research challenges; a confirmed duplication-positive clinical cohort is still missing.

Every verified file is listed

Complete public WES audit bundle

Includes indexed and unindexed manifests, summaries, the Dropbox audit, BRCA/SEQC2 inventories, range-test evidence, checksums, and reproducible scripts.

SHA-25670aa888624d6a82bd476978fa4f18fd5b3a1e55dc7dc7cdd29e7f2f9e3a61866

Official repositories and primary evidence

  1. International Genome Sample Resource. Official Phase 3 exome alignment release.
  2. European Nucleotide Archive. Advanced Search and Portal API.
  3. DepMap. Access to CCLE/DepMap genomic files.
  4. AWS Open Data. DepMap Cancer Cell Line Encyclopedia.
  5. NIST Genome in a Bottle. Official NCBI data repository.
  6. Baid G, et al. Gold-standard sequencing dataset for benchmarking.
  7. Fang LT, et al. SEQC2 cancer reference samples and call sets.
  8. SEQC2 Consortium. Somatic copy-number detection on HCC1395.
  9. NCI Genomic Data Commons. Open and controlled data access.