Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search
arXiv:2606.17209v1 Announce Type: new Abstract: Test-time scaling for agentic search typically increases depth (i.e., more turns and tokens per trajectory) or breadth (i.e., mor...
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
arXiv:2606.17220v1 Announce Type: new Abstract: Legal case retrieval remains challenging due to the complexity of legal language and the need for precise lexical alignment betwe...
SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions
arXiv:2606.17266v1 Announce Type: new Abstract: Production planning increasingly has to treat workforce capability as a decision variable: certifications lapse when skills are n...
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision
arXiv:2606.17162v1 Announce Type: new Abstract: Personalized presentation generation requires more than conditioning on a current prompt or template: agents must preserve stable...
Self-Generated Error Training for Token Editing in Diffusion Language Models
arXiv:2606.17175v1 Announce Type: new Abstract: Token-to-token (T2T) editing lets LLaDA2.1 revise committed tokens during block-diffusion decoding. The released recipe trains th...
Revisiting LLM Adaptation for 3D CT Report Generation: A Study of Scaling and Diagnostic Priors
arXiv:2606.17213v1 Announce Type: new Abstract: Recent advances in multimodal learning, including large language models (LLMs) and vision-language models (VLMs), have demonstrat...
Predicting model behavior before release by simulating deployment
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.
OpenAI frontier models and Codex are now available on AWS
OpenAI frontier models and Codex are now generally available on AWS, giving enterprises a new path to build with OpenAI through the AWS environments, controls, and procurement w...
Databricks brings GPT-5.5 to enterprise agent workflows
Databricks uses GPT-5.5 for enterprise agent workflows after the model set a new state of the art on the OfficeQA Pro benchmark.
CyberSecEval 2 - A Comprehensive Evaluation Framework for Cybersecurity Risks and Capabilities of Large Language Models
CyberSecEval 2 - A Comprehensive Evaluation Framework for Cybersecurity Risks and Capabilities of Large Language Models
Introducing the Gemini 2.5 Computer Use model
Available in preview via the API, our Computer Use model is a specialized model built on Gemini 2.5 Pro’s capabilities to power agents that can interact with user interfaces.
Which AI agent spent the money on your OpenAI/Anthropic bill
Article URL: https://github.com/Nu11P01nt3r3xc3pt10n/spaturzu-sdks Comments URL: https://news.ycombinator.com/item?id=48567889 Points: 1 # Comments: 0
A New Framework for Evaluating Voice Agents (EVA)
A New Framework for Evaluating Voice Agents (EVA)
Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic
Beyond LLMs: Why Scalable Enterprise AI Adoption Depends on Agent Logic
Gemma Scope 2: helping the AI safety community deepen understanding of complex language model behavior
Open interpretability tools for language models are now available across the entire Gemma 3 family with the release of Gemma Scope 2.
Strengthening our Frontier Safety Framework
We’re strengthening the Frontier Safety Framework (FSF) to help identify and mitigate severe risks from advanced AI models.
Ask HN: Does your mind drift while waiting for AI prompts to finish?
I've been a software engineer for 9 years now, and I noticed a very new weirdness in my workflows. Once I finish the architecture of a project and i have my context engineering...
Building an AI Agent in 6 Weeks (and Understanding How They Work)
Article URL: https://belderbos.dev/blog/jeff-haemer-agentic-ai-cohort/ Comments URL: https://news.ycombinator.com/item?id=48567620 Points: 1 # Comments: 0
Skill-Constrained Model Predictive Control for Resilient Manufacturing Supply Chains
arXiv:2606.17269v1 Announce Type: new Abstract: In skill-constrained production-inventory systems, the qualified human capacity available tomorrow depends on training decisions...
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty
arXiv:2606.17312v1 Announce Type: new Abstract: Large language models can arrive at the same answer through reasoning paths that are unstable, contradictory, or difficult to ran...