首页 > 专栏 > cs.CL updates on arXiv.org cs.CL updates on arXiv.org 共 2754 条资讯 The Authenticity Gap in Human Evaluation 2026-06-28 03:07:22 BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models 2026-06-28 03:07:22 On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification 2026-06-28 03:07:22 BayesPrompt: human readable prompts that make sense 2026-06-28 03:07:22 Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation 2026-06-28 03:07:22 Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints 2026-06-28 03:07:22 From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector 2026-06-28 03:07:22 Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE 2026-06-28 03:07:22 Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses 2026-06-28 03:07:22 FishBack: Pullback Fisher Geometry for Optimal Activation Steering in Transformers 2026-06-28 03:07:22 Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It 2026-06-28 03:07:22 Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks 2026-06-28 03:07:22 TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification 2026-06-28 03:07:22 Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review 2026-06-28 03:07:22 Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility 2026-06-28 03:07:22 The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT 2026-06-28 03:07:22 Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See 2026-06-28 03:07:22 Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges 2026-06-28 03:07:22 Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent 2026-06-28 03:07:22 Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback 2026-06-28 03:07:22 « 上一页1…56789…138下一页 » 相关分类 #!/slash/note #UNTAG (B)(F)uzzing on my world (Hi)story (IN)SECURE Magazine Notification (gdb) break *0x972 - 带鱼博客 BeltfishBlog - ./kwaa.dev .NET Blog .Trash /home/rook1e 00's Adventure 0kami's Blog 0x41414141 in ?? () 0x7f Blog 0xRick Owned Root ! 0xd00's blog 1 Byte 1A23 Blog 1A23 Studio 1Link.Fun 1stwebdesigner 251 2BAB 的工程博客 2ch中文网 360 CERT 360 Netlab Blog - Network Securi 38号车评中心 3o米的微博 404 Media