News & Current Events
Insights & Perspectives
AI Research
AI Products
Search
EDGENT
IQ
EDGENT
iq
About
Terms
Privacy Policy
Contact Us
EDGENT
iq
News & Current Events
Insights & Perspectives
Analytical Insights & Perspectives
Financial Sector Fortifies Against Surging AI-Powered Scams
Analytical Insights & Perspectives
Deloitte’s 2025 Outlook: Navigating Escalating AI Challenges in Human Capital
Analytical Insights & Perspectives
Salesforce Study Reveals Data Quality is Pivotal for Employee Trust in AI Adoption
Analytical Insights & Perspectives
Top Executives Sidestep Company AI Guidelines, Fueling Shadow AI Risks
Analytical Insights & Perspectives
Intel’s Evolving IP Strategy: A Calculated Shift Towards Core AI Innovation
Analytical Insights & Perspectives
Generative AI Prompts Increased Workforce Surveillance in Indian IT Sector
AI Research
AI Products
Search
EDGENT
IQ
News & Current Events
Insights & Perspectives
Analytical Insights & Perspectives
Financial Sector Fortifies Against Surging AI-Powered Scams
Analytical Insights & Perspectives
Deloitte’s 2025 Outlook: Navigating Escalating AI Challenges in Human Capital
Analytical Insights & Perspectives
Salesforce Study Reveals Data Quality is Pivotal for Employee Trust in AI Adoption
Analytical Insights & Perspectives
Top Executives Sidestep Company AI Guidelines, Fueling Shadow AI Risks
Analytical Insights & Perspectives
Intel’s Evolving IP Strategy: A Calculated Shift Towards Core AI Innovation
Analytical Insights & Perspectives
Generative AI Prompts Increased Workforce Surveillance in Indian IT Sector
AI Research
AI Products
Search
ProBench: A Deeper Look into How We Evaluate AI Agents for Mobile Apps
Google Unveils Free 5-Day AI Agents Intensive Course on Kaggle
A New Benchmark for Evaluating AI in Electronic Health Records: Introducing EHRStruct
FaithAct: A Framework for Verifying AI’s Visual Reasoning Steps
Evaluating AI in Journalism: A Practitioner-Centered Approach to Better Benchmarks
Recently Added
AI’s Hidden Costs: Gaps in Social Impact Reporting Revealed
Read more
AI’s Promise and Peril in Bangladeshi Law: A Dual Evaluation of Language Models
Read more
Evaluating Spatial Reasoning in Vision-Language Models with Referring Expressions
Read more
Evaluating AI’s Thought Process: A New Metric for Multimodal Reasoning
Read more
MVU-Eval: A New Benchmark for AI’s Multi-Video Understanding
Read more
PRIME: A New Framework to Diagnose AI’s Stereotypical Reasoning
Read more
Introducing Secu-Table: A New Dataset for AI to Understand Cybersecurity Data
Read more
Unpacking Construct Validity in Large Language Model Evaluations
Read more
IndicVisionBench: A New Frontier for Evaluating AI’s Cultural and Multilingual Understanding in India
Read more
Evaluating AI’s Coding Prowess in Ukrainian: Introducing UA-Code-Bench
Read more
AgileThinker: AI Agents Mastering Real-Time Decisions in Dynamic Environments
Read more
ChatGPT’s Persistent ‘Hallucinations’ Highlight AI Accuracy Challenges Despite Upgrades
Read more
New Benchmarking Suite Terminal-Bench 2.0 and Agent Testing Framework Harbor Launched to Advance AI Agent Evaluation
Read more
Laude Institute Unveils Inaugural ‘Slingshots’ AI Grants to Accelerate Innovation
Read more
Unpacking the Evolving Roles of AI: From Chatbots to Human-Centered Support Systems
Read more
Large Language Models as Judges for Recommender Systems: A New Approach to Understanding User Preferences
Read more
A Human-Centered Approach to Evaluating Voice AI Testing Platforms
Read more
RxSafeBench: A New Benchmark to Assess AI’s Medication Safety in Healthcare
Read more
Unveiling the Capabilities and Risks of the Jr. AI Scientist System
Read more
Databricks Unveils Advanced Evaluation Tools to Elevate AI Agent Performance and Governance
Read more
ChiMDQA: A New Comprehensive Dataset for Chinese Document Question Answering
Read more
AI Readiness Project Launched to Strengthen Public Sector’s Responsible AI Integration
Read more
OpenAI Introduces IndQA: A New Benchmark for AI Understanding of Indian Languages and Culture
Read more
Understanding LLM Reliability: A Hierarchical Imprecise Probability Framework
Read more
A New Arena for AI Coders: Benchmarking Goal-Oriented Software Development
Read more
Multi-Bench: A New Standard for Evaluating Emotional AI in Conversations
Read more
Unpacking AI’s Thought Process: A New Framework for Evaluating LLM Reasoning
Read more
Benchmarking LLMs for Cyber Threat Intelligence: Introducing AthenaBench
Read more
Speech-DRAME: A New Standard for Evaluating AI Speech Role-Play
Read more
Assessing LLM Defenses Against Prompt Injection: A New Evaluation Framework
Read more
Load more
Gen AI News and Updates
Subscribe
I have read and accepted the
Terms of Use
and
Privacy Policy
of the website and company.
- Advertisement -
What's new?
Search
ProBench: A Deeper Look into How We Evaluate AI Agents for Mobile Apps
November 14, 2025
Google Unveils Free 5-Day AI Agents Intensive Course on Kaggle
November 13, 2025
A New Benchmark for Evaluating AI in Electronic Health Records: Introducing EHRStruct
November 12, 2025
Load more