News & Current Events
Insights & Perspectives
AI Research
AI Products
Search
EDGENT
IQ
EDGENT
iq
About
Terms
Privacy Policy
Contact Us
EDGENT
iq
News & Current Events
Insights & Perspectives
Analytical Insights & Perspectives
Financial Sector Fortifies Against Surging AI-Powered Scams
Analytical Insights & Perspectives
Deloitte’s 2025 Outlook: Navigating Escalating AI Challenges in Human Capital
Analytical Insights & Perspectives
Salesforce Study Reveals Data Quality is Pivotal for Employee Trust in AI Adoption
Analytical Insights & Perspectives
Top Executives Sidestep Company AI Guidelines, Fueling Shadow AI Risks
Analytical Insights & Perspectives
Intel’s Evolving IP Strategy: A Calculated Shift Towards Core AI Innovation
Analytical Insights & Perspectives
Generative AI Prompts Increased Workforce Surveillance in Indian IT Sector
AI Research
AI Products
Search
EDGENT
IQ
News & Current Events
Insights & Perspectives
Analytical Insights & Perspectives
Financial Sector Fortifies Against Surging AI-Powered Scams
Analytical Insights & Perspectives
Deloitte’s 2025 Outlook: Navigating Escalating AI Challenges in Human Capital
Analytical Insights & Perspectives
Salesforce Study Reveals Data Quality is Pivotal for Employee Trust in AI Adoption
Analytical Insights & Perspectives
Top Executives Sidestep Company AI Guidelines, Fueling Shadow AI Risks
Analytical Insights & Perspectives
Intel’s Evolving IP Strategy: A Calculated Shift Towards Core AI Innovation
Analytical Insights & Perspectives
Generative AI Prompts Increased Workforce Surveillance in Indian IT Sector
AI Research
AI Products
Search
Standardizing Scientific Machine Learning: Introducing the MLCommons Benchmarks Ontology
New Benchmark Reveals AI Agents Struggle with Real-World Software Performance Optimization
Speech-DRAME: A New Standard for Evaluating AI Speech Role-Play
New Benchmarks Advance AI Mathematical Reasoning to Olympiad Levels
Salesforce AI Research Unveils WALT: A Novel Framework for Web Agents to Autonomously Discover and Utilize Website-Native Tools
Recently Added
IBM Research Unveils ‘Toucan,’ a Groundbreaking Dataset to Revolutionize AI Agent Tool-Calling Capabilities
Read more
BIGCODEARENA: Elevating Code Generation Evaluation Through Execution
Read more
Advancing Language Model Alignment: A New Approach to Long-Context Reward Modeling
Read more
BuilderBench: A New Frontier for Generalist AI Agents Learning Through Play
Read more
Assessing Value Consistency in Large Language Models with VAL-Bench
Read more
New Benchmark Reveals Language Model Vulnerabilities to Sociopolitical Harms
Read more
The Evolving Landscape of AI Evaluation: From Simple Recognition to Complex Reasoning
Read more
OptunaHub: Centralizing Black-Box Optimization for Enhanced Research
Read more
Unpacking AI’s Ethical Compass: How LLMs Allocate Social Welfare
Read more
TRACE: A Framework for Dynamically Evolving AI Agent Benchmarks
Read more
Measuring Personalization in Deep Research Agents
Read more
MULocBench: A New Benchmark for Pinpointing Software Issues Beyond Code
Read more
DOoM Benchmark: Challenging Language Models with Russian Math and Physics Problems
Read more
Alibaba’s Qwen Team Unveils Qwen3-Coder: A 480 Billion Parameter Open-Source AI for Advanced Coding
Read more
Evaluating LLM Precision: A New Benchmark for Function Calling Instruction Adherence
Read more
Unlocking Auditory Imagination in Language Models: The AuditoryBench++ Benchmark and AIR-CoT Method
Read more
Beyond the Hype: Assessing GPT-5’s Impact and the Critical Need for AI Security in Engineering
Read more
EdiVal-Agent: A New Standard for Evaluating Multi-Turn Image Editing
Read more
Critical Infrastructure’s AI Reckoning: Nuclear Sector Benchmarks End Self-Governance Era, Demand Policy & Ethics Intervention
Read more
The Hidden Vulnerability of LLMs: Performance Drops with Reworded Questions
Read more
xAI Unveils ‘Grok Code Fast 1’: A New Speedy and Economical AI Agent for Developers
Read more
Evaluating AI on Unanswered Questions: A New Benchmark for Language Models
Read more
Measuring AI Agent Performance and Safety in Online Shopping
Read more
Benchmarking AI’s Tool-Using Abilities in the Real World
Read more
LLaSO: An Open Standard for Large Speech-Language Models
Read more
Kitchen-R: A New Benchmark for Integrated Robot Planning and Control in Simulated Kitchens
Read more
Boosting LLM Reasoning: A Deep Dive into Tool Integration and Efficiency
Read more
VisCodex: A Unified System for Generating Code from Images
Read more
New Benchmark Reveals Language Models Struggle with Video Game Logic and Spatial Reasoning
Read more
Klear-Reasoner: Enhancing AI Reasoning Through Refined Training Methods
Read more
Load more
Gen AI News and Updates
Subscribe
I have read and accepted the
Terms of Use
and
Privacy Policy
of the website and company.
- Advertisement -
What's new?
Search
Standardizing Scientific Machine Learning: Introducing the MLCommons Benchmarks Ontology
November 11, 2025
New Benchmark Reveals AI Agents Struggle with Real-World Software Performance Optimization
November 11, 2025
Speech-DRAME: A New Standard for Evaluating AI Speech Role-Play
November 4, 2025
Load more