Glossary · Multimodal systems

Early Fusion

Combining raw or low-level representations from several modalities before most task-specific modeling occurs.

Why it matters

Early interaction can expose fine-grained cross-modal relationships, but it also requires compatible representations and careful handling of alignment and missing inputs.

In practice

Convert each modality into a declared token or feature representation, preserve source and position markers, fuse them before the shared backbone, and compare against single-modality and late-fusion baselines.

Common confusion

Early fusion describes where streams are combined in the architecture. It does not guarantee that the model learns useful alignment between them.

Related terms

Sources

Browse the learning paths to see this term in context — every lesson is free to read.