Marine DebrisCompetition research · v7

ENSEMBLES · REPRODUCIBLE FALLBACK · TEAM RESULT HANDOFF

Ensemble Evolution

The current work is not broad fusion search. It is reconstruction of the strongest team pipeline, measured replacement tests, and retaining only components that demonstrably help.

TEAM HIDDEN ≈0.7701REPRODUCIBLE HIDDEN 0.7587924RESCORER PROVENHANDOFF IN PROGRESS
≈0.7701best observed team hidden
0.7587924best fully reproducible hidden
0.746837507paired reproducible local mAP50
0.748909485three-model local candidate
Best fully reproducible fallback system
FieldControlled value
ModelsCo-DINO 1280 scale-control + MI-DETR 1024
MethodClass-aware Weighted Box Fusion
Confidence aggregationmax
IoU / skip0.45 / 0.001
Model weights1.0 / 1.0
Outputstable top-300/image · PURE_INTEGER_XYWH submission
Local canonical / AIdea hidden0.7468375071854423 / 0.7587924
StatusFULLY REPRODUCIBLE FALLBACK
STRONGEST CONTROLLED SYSTEM

Co-DINO1280 + MI-DETR1024

The complete local provenance, prediction artifacts, and equal-weight max-confidence WBF recipe are reconstructed and controlled.

TEAM UPSIDE

Crop-rescored pipeline

The stronger ≈0.7701 hidden result is credible competition evidence, but it remains reproduction-pending until its complete handoff is reconstructed locally.

Ensemble development record · local and hidden are separate
MethodLocal mAP50AIdea hiddenDecision / interpretation
Co-DINO1280 + MI-DETR1024 · equal-weight WBF0.74683750720.7587924FULLY REPRODUCIBLE FALLBACK
Co-DINO1280 + RT-DETRv4-X12800.74143619710.7415787REJECTED AS PRIMARY TWO-MODEL FUSION
Co-DINO1280 + MI-DETR1024 + RT-DETRv4-X12800.7489094849LOCALLY SUPPORTED THIRD-MEMBER CANDIDATE
Teammate ensemble + ConvNeXt-Tiny rescorerapproximately 0.7701BEST OBSERVED · REPRODUCTION HANDOFF PENDING
Prior 1024 equal-weight WBF0.74374010610.7534597HISTORICAL REFERENCE
Class-bucket WBF0.74458629660.7534510REJECTED · NOT CANONICAL
Tiny-area gated RT-DETRv40.74167042230.7484452REJECTED · HIDDEN REGRESSION
RECALL CONTRIBUTION

RT-DETRv4 as third member

The three-model candidate adds approximately 389 unique TP over the two-model base, around 201 under 1% object area and around 28 under 0.1%. It is evidence-supported, not automatically retained.

WHY TWO-MODEL RTv4 FAILED

Quality burden

With Co-DINO alone, RT-DETRv4 adds recall but also too much lower-quality prediction burden. Its role must be tested within the stronger rescorer pipeline, not assumed.

NO BLIND ADDITIONS

Measure replacement value

A model remains only if it helps the reconstructed final system, not merely because it has a good standalone or local-fusion number.

Current ensemble system path

Complete strongest Pareto variantsCo-DINO and MI-DETR priority; RT-DETRv4 remains an additive candidate
Reconstruct teammate ≈0.7701 pipelineCollect exact artifacts, recipe, preprocessing, scoring, and submission provenance
Freeze strongest Co-DINO + MI-DETR pairReplace individual inputs only through measured local and hidden evidence
Retain RT-DETRv4 only if additiveEspecially for tiny/crowded-object recovery
Final preliminary submissionUse the fewest justified components