Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- View PDF HTML (experimental) Abstract:Multi-LLM systems harness the complementary strengths of diverse Large Language Models, achieving performance and efficiency gains that are not attainable by a single model.
- In existing designs, LLMs communicate through text, forcing internal representations to be transformed into output token sequences.
- This process both loses rich semantic information and incurs token-by-token generation latency.
Unverified
- View PDF HTML (experimental) Abstract:Multi-LLM systems harness the complementary strengths of diverse Large Language Models, achieving performance and efficiency gains that are not attainable by a single model.
- In existing designs, LLMs communicate through text, forcing internal representations to be transformed into output token sequences.
- This process both loses rich semantic information and incurs token-by-token generation latency.
Sources: Arxiv