Cross-Modal Representation Learning for Generative Intelligence Systems

Authors

  • Ivan Novak Research Scientist, Institute of Intelligent Systems, Swiss Institute of Machine Intelligence, Zurich, Switzerland Author
  • Hugo Horvath Assistant Professor, Department of Computer Science, Nordic Technical University, Stockholm, Sweden Author
  • Eva Garcia Research Scientist, Department of Computer Science, Central European Tech University, Vienna, Austria Author

Keywords:

cross-modal representation learning, contrastive learning, multimodal, hierarchical alignment, compositional generalisation, CLIP, generative intelligence, semantic alignment

Abstract

Cross-modal representation learning -- the development of shared latent representations that capture semantic correspondences across different data modalities -- is a foundational capability for generative intelligence systems that must reason and generate coherently across text, image, audio, and other modality spaces. Despite substantial progress in bimodal alignment (CLIP for text-image, CLAP for text-audio) and emerging trimodal approaches, the theoretical foundations and systematic evaluation of cross-modal representation quality remain underdeveloped relative to the engineering advances. This paper proposes the Cross-Modal Representation Quality (CMRQ) framework, a theoretical and empirical framework for evaluating the quality of cross-modal representations along five dimensions: semantic alignment, structural preservation, compositional generalization, cross-modal transfer efficiency, and generation coherence. The CMRQ framework is instantiated through a novel Hierarchical Cross-Modal Contrastive (HCMC) learning objective that extends standard contrastive alignment to three hierarchical levels of semantic granularity -- instance, category, and attribute -- enabling richer, more compositionally structured cross-modal representations. HCMC is evaluated on eight cross-modal benchmarks spanning text-image, text-audio, and image-audio modality pairs, and compared against CLIP, CLAP, and ImageBind baselines. HCMC achieves state-of-the-art cross-modal retrieval performance on six of eight benchmarks, with mean R@1 improvement of 8.4 percentage points over CLIP and 11.2 points over ImageBind. Compositional generalisation performance -- evaluated on novel attribute-object combinations not seen during training -- improves by 24.6% over CLIP, demonstrating that hierarchical contrastive learning produces structurally richer cross-modal representations. The study contributes the CMRQ evaluation framework, the HCMC learning objective, and a compositional cross-modal evaluation benchmark.

Author Biographies

  • Ivan Novak, Research Scientist, Institute of Intelligent Systems, Swiss Institute of Machine Intelligence, Zurich, Switzerland

    Research Scientist, Institute of Intelligent Systems, Swiss Institute of Machine Intelligence, Zurich, Switzerland

  • Hugo Horvath, Assistant Professor, Department of Computer Science, Nordic Technical University, Stockholm, Sweden

    Assistant Professor, Department of Computer Science, Nordic Technical University, Stockholm, Sweden

  • Eva Garcia, Research Scientist, Department of Computer Science, Central European Tech University, Vienna, Austria

    Research Scientist, Department of Computer Science, Central European Tech University, Vienna, Austria

Downloads

Published

2024-03-28

How to Cite

Cross-Modal Representation Learning for Generative Intelligence Systems. (2024). Journal of Generative Intelligence E: 3117-6429 P: 3117-6437, 1(1), 41-48. https://galaxiauniverse.com/index.php/JGI/article/view/289