Glossary · Multimodal systems
Early Fusion
Combining raw or low-level representations from several modalities before most task-specific modeling occurs.
Why it matters
Early interaction can expose fine-grained cross-modal relationships, but it also requires compatible representations and careful handling of alignment and missing inputs.
In practice
Convert each modality into a declared token or feature representation, preserve source and position markers, fuse them before the shared backbone, and compare against single-modality and late-fusion baselines.
Common confusion
Early fusion describes where streams are combined in the architecture. It does not guarantee that the model learns useful alignment between them.
Related terms
Sources
Browse the learning paths to see this term in context — every lesson is free to read.