| # |
Team |
Task |
Date/Time |
DataID |
AMFM |
Method
|
Other Resources
|
System Description |
| unuse |
unuse |
unuse |
unuse |
unuse |
unuse |
unuse |
unuse |
unuse |
unuse |
|
| 1 | CNLP-NITS-PP | MMEVMM24en-bn | 2022/07/11 12:51:09 | 6743 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | Transliteration-based phrase pairs augmentation and visual features in training using BRNN encoder and doubly-attentive-rnn decoder. |
| 2 | SILO_NLP | MMEVMM24en-bn | 2022/07/14 20:56:54 | 6939 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | Object Tags (Image) + Finetune mBART |
| 3 | ODIAGEN | MMEVMM24en-bn | 2023/07/06 04:02:32 | 7107 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | Image features extracted as Object tags appended with text and MBART fine-tuning |
| 4 | BITS-P | MMEVMM24en-bn | 2023/07/08 13:40:55 | 7123 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | Yes | NLLB model finetuned on captions + object tags of original & synthetic images using DETR model |
| 5 | 1128 | MMEVMM24en-bn | 2024/07/23 15:43:17 | 7137 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | initial model |
| 6 | 00-7 | MMEVMM24en-bn | 2024/08/05 15:15:26 | 7191 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | Yes | TEST |
| 7 | v036 | MMEVMM24en-bn | 2024/08/11 12:34:14 | 7317 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | NMT based system using both image descriptors and text description. A multistage LLM pipeline used for extracting image data descriptions and translation. Fine tuning done in few cases
Models Used:
|
| 8 | 239233 | MMEVMM24en-bn | 2024/08/13 12:28:24 | 7381 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | Yes | One-shot prompt for synthetic QA description from captions; translate QA using IndicTrans2; generate caption from QA as context |
| 9 | UNLP | MMEVMM24en-bn | 2024/08/13 17:55:18 | 7391 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | Using the Transformer-based Gated Fusion model to integrate both text and visual data. |
| 10 | v036 | MMEVMM24en-bn | 2024/08/15 19:53:41 | 7417 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | |
| 11 | v036 | MMEVMM24en-bn | 2024/08/15 19:55:30 | 7418 | - | - | - | - | - | - | 0.000000 | - | - | - | NMT | No | |
| 12 | IITP-AI-NLP-ML | MMEVMM24en-bn | 2025/10/22 21:33:19 | 7457 | - | - | - | - | - | - | 0.000000 | - | - | - | SMT | Yes | Used Selective Attention Architecture with IndicTrans as the base model and CLIP ViT-B/16 model to extract image features. We extract a) Full image feats, and b) cropped image feats and pick the one w |