Список литературы
1. Hendrycks D., Burns C., Basart S. [et al.]. Measuring Massive Multitask Language Understanding [Электронный ресурс] // arXiv. 2021. arXiv: 2009.3 300. DOI: 10.48 550/arXiv.2009.3 300. URL:
arxiv.org/abs/2009.3 300 (дата обращения: 16.04.2026).
2. Fenogenova A., Chervyakov A., Martynov N. [et al.]. MERA: A Comprehensive LLM Evaluation in Russian // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok: Association for Computational Linguistics, 2024. P. 9920−9948. DOI: 10.18 653/v1/2024.acl-long.534.
3. Shavrina T., Fenogenova A., Emelyanov A. [et al.]. RussianSuperGLUE: A Russian Language Understanding Evaluation Benchmark // Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Stroudsburg: Association for Computational Linguistics, 2020. P. 4717−4726. DOI: 10.18 653/v1/2020.emnlp-main.381.
4. Chiang W.-L., Zheng L., Sheng Y. [et al.]. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference // Proceedings of the 41st International Conference on Machine Learning (ICML). Vienna: PMLR, 2024. P. 8359−8388.
5. Zheng L., Chiang W.-L., Sheng Y. [et al.]. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena // Advances in Neural Information Processing Systems (NeurIPS). 2023. Vol. 36. P. 46 595−46 623. DOI: 10.48 550/arXiv.2306.5 685.
6. Zhu L., Wang X., Wang X. JudgeLM: Fine-tuned Large Language Models are Scalable Judges // Proceedings of the 13th International Conference on Learning Representations (ICLR). 2025. DOI: 10.48 550/arXiv.2310.17 631.
7. Qwen Team. Qwen3 Technical Report [Электронный ресурс] // arXiv. 2025. arXiv: 2505.9 388. DOI: 10.48 550/arXiv.2505.9 388. URL:
arxiv.org/abs/2505.9 388 (дата обращения: 22.04.2026).
8. Yang A., Yang B., Hui B. [et al.]. Qwen2 Technical Report [Электронный ресурс] // arXiv. 2024. arXiv: 2407.10 671. DOI: 10.48 550/arXiv.2407.10 671. URL:
arxiv.org/abs/2407.10 671 (дата обращения: 25.04.2026).
9. Google DeepMind. Gemma 3 Technical Report [Электронный ресурс] // arXiv. 2025. arXiv: 2503.19 786. DOI: 10.48 550/arXiv.2503.19 786. URL:
arxiv.org/abs/2503.19 786 (дата обращения: 28.04.2026).
10. Abdin M., Aneja J., Behl H. [et al.]. Phi-4 Technical Report [Электронный ресурс] // arXiv. 2024. arXiv: 2412.8 905. DOI: 10.48 550/arXiv.2412.8 905. URL:
arxiv.org/abs/2412.8 905 (дата обращения: 30.04.2026).
11. Jiang A. Q., Sablayrolles A., Roux A. [et al.]. Mixtral of Experts [Электронный ресурс] // arXiv. 2024. arXiv: 2401.4 088. DOI: 10.48 550/arXiv.2401.4 088. URL:
arxiv.org/abs/2401.4 088 (дата обращения: 04.05.2026).
12. Adler B., Agarwal N., Aithal A. [et al.]. Nemotron-4 340B Technical Report [Электронный ресурс] // arXiv. 2024. arXiv: 2406.11 704. DOI: 10.48 550/arXiv.2406.11 704. URL:
arxiv.org/abs/2406.11 704 (дата обращения: 06.05.2026).
13. GLM Team. ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools [Электронный ресурс] // arXiv. 2024. arXiv: 2406.12 793. DOI: 10.48 550/arXiv.2406.12 793. URL:
arxiv.org/abs/2406.12 793 (дата обращения: 08.05.2026).
14. Groeneveld D., Beltagy I., Walsh P. [et al.]. OLMo: Accelerating the Science of Language Models // Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Bangkok: Association for Computational Linguistics, 2024. P. 15 789−15 809. DOI: 10.18 653/v1/2024.acl-long.841.
15. Gao L., Tow J., Abbasi B. [et al.]. A Framework for Few-Shot Language Model Evaluation [Электронный ресурс]. Zenodo, 2023. DOI: 10.5281/zenodo.10 256 836. URL:
doi.org/10.5281/zenodo.10 256 836 (дата обращения: 12.05.2026).
16. Kadavath S., Conerly T., Askell A. [et al.]. Language Models (Mostly) Know What They Know [Электронный ресурс] // arXiv. 2022. arXiv: 2207.5 221. DOI: 10.48 550/arXiv.2207.5 221. URL:
arxiv.org/abs/2207.5 221 (дата обращения: 14.05.2026).
17. Qwen Team. Qwen3.5: Towards Native Multimodal Agents [Электронный ресурс]. 2026. URL:
qwen.ai/blog?id=qwen3.5 (дата обращения: 02.06.2026).
18. Google DeepMind. Gemma 4: The most capable open models [Электронный ресурс]. 2026. URL:
blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ (дата обращения: 03.06.2026).
19. Zhipu AI (Z.ai). GLM-4.7-Flash: An efficient MoE model for local coding and agents [Электронный ресурс]. 2026. URL:
www.zhipuai.cn/en/news/148 (дата обращения: 04.06.2026).
20. OpenAI. gpt-oss-120b & gpt-oss-20b Model Card [Электронный ресурс] // arXiv. 2025. arXiv: 2508.10 925. DOI: 10.48 550/arXiv.2508.10 925. URL:
arxiv.org/abs/2508.10 925 (дата обращения: 05.06.2026).
21. DeepSeek-AI. DeepSeek-V4 Technical Report [Электронный ресурс]. 2026. URL:
huggingface.co/deepseek-ai/DeepSeek-V4-Pro (дата обращения: 06.06.2026).