首页 > 专栏 > cs.CL updates on arXiv.org cs.CL updates on arXiv.org 共 2661 条资讯 When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling 2026-06-28 03:07:22 SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces 2026-06-28 03:07:22 ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators 2026-06-28 03:07:22 Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimation 2026-06-28 03:07:22 Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework 2026-06-28 03:07:22 Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism 2026-06-28 03:07:22 DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents 2026-06-28 03:07:22 GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation 2026-06-28 03:07:22 Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale 2026-06-28 03:07:22 An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms 2026-06-28 03:07:22 Parameter Golf: What Really Works? 2026-06-28 03:07:22 Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents 2026-06-28 03:07:22 From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages 2026-06-28 03:07:22 Comparing Architectures for Supervised Political Scaling 2026-06-28 03:07:22 eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian 2026-06-28 03:07:22 Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting 2026-06-28 03:07:22 Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation 2026-06-28 03:07:22 FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning 2026-06-28 03:07:22 StatEval: A Comprehensive Benchmark for Large Language Models in Statistics 2026-06-28 03:07:22 IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs 2026-06-28 03:07:22 « 上一页1…110111112113114…134下一页 » 相关分类 #!/slash/note #UNTAG (B)(F)uzzing on my world (Hi)story (IN)SECURE Magazine Notification (gdb) break *0x972 - 带鱼博客 BeltfishBlog - ./kwaa.dev .NET Blog .Trash /home/rook1e 00's Adventure 0kami's Blog 0x41414141 in ?? () 0x7f Blog 0xRick Owned Root ! 0xd00's blog 1 Byte 1A23 Blog 1A23 Studio 1Link.Fun 1stwebdesigner 251 2BAB 的工程博客 2ch中文网 360 CERT 360 Netlab Blog - Network Securi 38号车评中心 3o米的微博 404 Media