Stratiformer: Depth-Aware Heterogeneous Layer Design for Efficient Language Models
Allocating model capacity heterogeneously across depth for more efficient language models.
Exploratory ideas, early-stage projects, and technical notes. For peer-reviewed work and technical reports, see my publications.
Allocating model capacity heterogeneously across depth for more efficient language models.
Agents that decide when and how to scale inference-time compute.
Competition-driven reasoning for selecting and refining candidate solutions.