-
Figure 1.
Performance comparison between the Bash/AWK reference pipeline and hickit_rs on an eight-core 32-GB workstation. (a) Processing time for 100 million and 1 billion read pairs on a logarithmic scale; hickit_rs completes both scales in approximately 30 s and 5 min, respectively, whereas the Bash/AWK pipeline takes roughly 250 s and 3,000 s—an 8.3–10× speedup. (b) Peak memory consumption. The Bash/AWK pipeline peaks at approximately 2,400 MB, whereas hickit_rs uses roughly 300–350 MB, an 86%–88% reduction (7–8× lower). Benchmarks were performed on an eight-core (Intel i7-10700 @ 2.9 GHz) 32-GB RAM workstation running Ubuntu 22.04. Input data were derived from the publicly available human Hi-C dataset SRR1658573 (NCBI SRA), processed to 100 million and 1 billion valid read pairs in Juicer merged_nodups format (gzip-compressed). hickit_rs v0.4 was invoked with the default parameters (hickit resolution merged_nodups.txt.gz); the reference pipeline used Juicer's AWK-based calculate_map_resolution.sh. Wall-clock time and peak memory were measured using /usr/bin/time -v; each benchmark was run three times with the median reported.
-
Figure 2.
A tiered engineering framework for AI-assisted development with four components: ① Specification-driven design (SDD), ② test-driven development (TDD), ③ multimodel exploration, and ④ periodic review. These are applied sequentially, with depth scaled to the task's complexity (bottom bar). Each component's primary defence against the three bioinformatics-specific failure modes discussed in the main text is annotated within its panel.
Figures
(2)
Tables
(0)