This dataset documents FlexiDepth's layer allocation patterns using Llama-3-8B-Instruct as the base model, as described in the paper Adaptive Layer-skipping in Pre-trained LLMs. It captures layer usage per token across two domains: language comprehension and math reasoning, revealing how FlexiDepth dynamically adjusts its depth based on task complexity.
Text Generation: The dataset includes 100 paragraphs randomly sampled from the XSum test set and evaluates three subtasks: copying… See the full description on the dataset page:
https://huggingface.co/datasets/xuan-luo/FlexiPatterns-Llama-3-8B-Instruct.