- Authors
- Michaela Dobrovolná, Georgie Middleton, Stefan Bidula, Vratislav Peška, et al.
- Published in
- BMC Plant Biology · 2026-06 · Institute of Biophysics of the Czech Academy of Sciences
- ProteinIQ in this study
- ProteinIQ's random DNA generator supplied 100 control sequences for comparisons with chloroplast genomes across ten groups.
- Tools used
- Random Dna
Finding a pattern in DNA is only the beginning. A useful next question is whether comparable random sequences would produce the same result.
The challenge: a fair comparison
Researchers studying 7,187 chloroplast genomes needed a baseline for interpreting predicted G-quadruplex-forming sequences, DNA motifs that may fold into four-stranded structures. Their comparison needed to account for both sequence length and GC content, the proportion of bases that are G or C.
An arbitrary random sequence would make a poor control. A longer sequence offers more opportunities to find a motif, while a different base composition changes the material from which those motifs can form.
The solution: generate controls in the browser
The team used ProteinIQ's DNAGenIQ random DNA generator to produce ten controls per group, 100 in total. Controls matched each group's average length and closely approximated its average GC content. G4Hunter then analysed the controls and chloroplast DNA with the same settings.
DNAGenIQ makes sequence count, length, and target GC content available as form settings. It runs in the browser and returns FASTA sequences, so this preparation step needs neither a local software installation nor a custom sequence-generation script.
For a researcher assembling a similar comparison, the practical benefit is having the input and output steps in one place. Set the required sequence properties, generate the records, and download a FASTA file for the next analysis. Each record's header includes its length and measured GC percentage, making it possible to inspect the generated controls before using them.
The target GC percentage controls random sampling; individual sequences can vary around it. Saving the generated file also matters: a fresh run produces new sequences, so retaining the original controls is necessary when repeating a comparison.
The outcome: evidence beyond base composition
The controls generally contained fewer predicted motifs than chloroplast DNA. The paper reports a 72% lower frequency in the 1.2 to 1.4 G4Hunter score interval, averaged across sequences. The heterogeneous "Other" category was an exception to the broader enrichment pattern.
The comparison supported a role for sequence organisation beyond overall GC content. Whether these predicted structures have biological functions still requires experimental validation. Read the published study.
For similar projects, DNAGenIQ provides a ready-to-use way to prepare random controls. Researchers still choose the comparison design and interpret the downstream analysis, with no need to build a generator just to obtain its inputs.



