Cross-Modal Representation Learning for Generative Intelligence Systems
Keywords:
cross-modal representation learning, contrastive learning, multimodal, hierarchical alignment, compositional generalisation, CLIP, generative intelligence, semantic alignmentAbstract
Cross-modal representation learning -- the development of shared latent representations that capture semantic correspondences across different data modalities -- is a foundational capability for generative intelligence systems that must reason and generate coherently across text, image, audio, and other modality spaces. Despite substantial progress in bimodal alignment (CLIP for text-image, CLAP for text-audio) and emerging trimodal approaches, the theoretical foundations and systematic evaluation of cross-modal representation quality remain underdeveloped relative to the engineering advances. This paper proposes the Cross-Modal Representation Quality (CMRQ) framework, a theoretical and empirical framework for evaluating the quality of cross-modal representations along five dimensions: semantic alignment, structural preservation, compositional generalization, cross-modal transfer efficiency, and generation coherence. The CMRQ framework is instantiated through a novel Hierarchical Cross-Modal Contrastive (HCMC) learning objective that extends standard contrastive alignment to three hierarchical levels of semantic granularity -- instance, category, and attribute -- enabling richer, more compositionally structured cross-modal representations. HCMC is evaluated on eight cross-modal benchmarks spanning text-image, text-audio, and image-audio modality pairs, and compared against CLIP, CLAP, and ImageBind baselines. HCMC achieves state-of-the-art cross-modal retrieval performance on six of eight benchmarks, with mean R@1 improvement of 8.4 percentage points over CLIP and 11.2 points over ImageBind. Compositional generalisation performance -- evaluated on novel attribute-object combinations not seen during training -- improves by 24.6% over CLIP, demonstrating that hierarchical contrastive learning produces structurally richer cross-modal representations. The study contributes the CMRQ evaluation framework, the HCMC learning objective, and a compositional cross-modal evaluation benchmark.
