About Sci2Pol
Writing a policy brief requires selecting the findings that matter for policy, explaining them clearly, and checking that recommendations follow from the evidence. Sci2Pol studies how well language models perform this work and whether targeted fine-tuning improves their performance.
We introduce Sci2Pol-Bench, with 18 tasks spanning autocompletion, understanding, summarization, generation, and verification, and evaluate 13 models. For brief generation, we develop an LLM-based evaluation metric aligned with expert judgments after finding that ROUGE and BERTScore miss important differences in writing quality.
We also construct Sci2Pol-Corpus by linking scientific papers to policy documents, filtering the resulting candidates, and revising the selected briefs with expert-written examples as references. The corpus contains 639 paper–brief pairs. Fine-tuning LLaMA-3.1-8B, Gemma-12B, and Gemma-27B improves performance across Sci2Pol-Bench; the fine-tuned Gemma-27B exceeds GPT-4o and DeepSeek-V3 on this benchmark.
Authors: Weimin Wu, Alexander Furnas, Eddie Yang, Gefei Liu, Akhil Pandey Akella, Xuefeng Song, Dashun Wang, Han Liu
Paper: ICLR 2026 proceedings
Code: Sci2Pol
HuggingFace: Sci2Pol-Bench · Sci2Pol-Corpus
