FreeKnob Audit — Calibration vs Predictive Skill
A model-agnostic audit for threshold-based rare-event evaluation. Matched monotone calibration separates predictive skill from threshold placement without retraining or changing spatial order.
Reliable Machine Learning · Uncertainty Quantification · Trustworthy Evaluation
I study how calibration and evaluation choices affect model comparisons, and develop uncertainty methods for deployment.
Bangladesh University of Engineering and Technology (BUET)
CGPA 3.62/4.00 · Undergraduate thesis: ShasthoSheba · Supervisor: Prof. A. B. M. Alim Al Islam · GRE 329 (Q170, V159) · IELTS 8.0
Work on reliable evaluation, uncertainty quantification, geospatial ML, and earlier human-centred computing.
ICLR 2027 submission · Datasets & Benchmarks
A matched-calibration audit for thresholded rare-event benchmarks. On released CasCast checkpoints, the cascade-over-backbone CSI gap falls from 0.1601 to 0.0339 after monotone recalibration; across 450 SEVIR contrasts, 51 reverse sign.
Manuscript
A post-hoc conformal wrapper for frozen crowd counters. Across four counters and four datasets, crop-consistency improves hardest-region coverage in every evaluated model–dataset pair; at matched coverage, intervals are 15% narrower than a label-trained uncertainty model and 29% narrower than isotonic recalibration.
ACM Transactions on Computer-Human Interaction · Under review
Heliyon, 10(1), e23100
Manuscript in preparation
Exact certificates for whether forecast rankings survive every acceptable calibration-improving monotone correction.
Climate Services User Forum, South Asian Hydromet Forum (SAHF) · Malé, Maldives
On behalf of RIMES
Capacity Building Training, National Center of Meteorology · UAE
On behalf of RIMES
Selected empirical research, applied research, and research software.
A model-agnostic audit for threshold-based rare-event evaluation. Matched monotone calibration separates predictive skill from threshold placement without retraining or changing spatial order.
A post-hoc conformal wrapper for frozen crowd counters using a label-free crop-consistency difficulty signal. Evaluated across four counters and four datasets, including cross-dataset weighted conformal prediction.
An exact rational-arithmetic framework for testing whether a fixed-threshold forecast ranking survives every acceptable monotone recalibration. Manuscript in preparation.
Ongoing research on when conformal guarantees break under vocabulary change and how the failure can be predicted from calibration data before deployment.
Deployed infrastructure, applied ML systems, and open-source contributions.
At AI GeoLAB, a QLoRA-tuned Qwen2.5-VL-3B reaches 0.965 core-field F1 on 10,509 held-out records. A separate segmentation and number-reading pipeline recovers and correctly numbers 91.7% of plots on 22 held-out cadastral sheets.
Led backend, data-pipeline, and ML architecture for a national meteorological platform processing more than 1 TB/day from five NWP models and two satellites. The system also supports operational CAP v1.2 alerting.
AI GeoLAB Ltd
Developing evaluated, human-in-the-loop methods for handwritten Bengali land records and cadastral map digitisation.
Regional Integrated Multi-Hazard Early Warning System (RIMES)
Led backend, data-pipeline, and ML architecture for Timor-Leste CDIS; designed CAP v1.2 alerting and the data unification behind Bangladesh FFWC's live 15-day flood forecasts.
Interactive Cares
Led an eight-person engineering team as the learning platform grew from 8,000 to 100,000 users.