We introduce VectorGym, a new comprehensive multi-task benchmark for evaluating Vision-Language Models (VLMs) on Scalable Vector …
Juan A. Rodriguez,
Haotian Zhang, Abhay Puri,
Tianyang Zhang,
Rishav Pramanik,
Meng Lin,
Xiaoqing Xie,
Marco Terral,
Darsh Kaushik,
Aly Shariff,
Perouz Taslakian, Spandana Gella,
Sai Rajeswar Mudumba,
David Vazquez, Christopher Pal,
Marco Pedersoli
Conference on Empirical Methods in Natural Language Processing (EMNLP),
2026.
Building reliable computer-use agents requires grounding: accurately connecting natural language instructions to the correct on-screen …
Aarash Feizi,
Shravan Nayak,
Xiangru Jian,
Kevin Qinghong Lin,
Kaixin Li,
Rabiul Awal, Xing Han Lu,
Johan Obando,
Juan A. Rodriguez,
Nicolas Chapados,
David Vazquez,
Adriana Romero Soriano,
Reihaneh Rabbany,
Perouz Taslakian, Christopher Pal, Spandana Gella,
Sai Rajeswar Mudumba
International Conference on Learning Representations,
2026.
Workflows are a fundamental component of automation in enterprise platforms, enabling the orchestration of tasks, data processing, and …
Patrice Béchard,
Chao Wang,
Juan A. Rodriguez,
Amirhossein Abaskohi, Christopher Pal,
David Vazquez, Spandana Gella,
Sai Rajeswar Mudumba,
Perouz Taslakian
European Chapter of the Association for Computational Linguistics (EACL),
2026.
Aligning visual features with language embeddings is a key challenge in vision-language models (VLMs). The performance of such models …
Ahmed Masry,
Juan A. Rodriguez,
Tianyu Zhang,
Suyuchen Wang,
Chao Wang,
Aarash Feizi,
Akshay Kalkunte, Abhay Puri,
Xiangru Jian, Pierre-André Noël,
Sathwik Madhusudhan,
Marco Pedersoli,
Bang Liu,
Nicolas Chapados,
Yoshua Bengio,
Enamul Hoque Prince , Christopher Pal, Issam H. Laradji,
David Vazquez,
Perouz Taslakian, Spandana Gella,
Sai Rajeswar Mudumba
Neural Information Processing Systems (NeurIPS),
2025.
Scalable Vector Graphics (SVG) offer a powerful format for representing visual designs as interpretable code. Recent advances in …
Juan A. Rodriguez,
Haotian Zhang, Abhay Puri,
Rishav Pramanik,
Aarash Feizi,
Pascal Wichmann,
Arnab Mondal,
Mohammad Reza Samsami,
Rabiul Awal,
Perouz Taslakian, Spandana Gella,
Sai Rajeswar Mudumba,
David Vazquez, Christopher Pal,
Marco Pedersoli
Neural Information Processing Systems (NeurIPS),
2025.
While image generation techniques are now capable of producing high quality images that respect prompts which span multiple sentences, …
Saba Ahmadi,
Rabiul Awal,
Ankur Sikarwar,
Amirhossein Kazemnejad,
Ge Ya Luo,
Juan A. Rodriguez,
Sai Rajeswar Mudumba, Siva Reddy, Christopher Pal,
Benno Krojer,
Aishwarya Agrawal
Neural Information Processing Systems (NeurIPS),
2025.