spot_img
HomeResearch & DevelopmentVideoConviction: A New Benchmark for Understanding Finfluencer Impact on...

VideoConviction: A New Benchmark for Understanding Finfluencer Impact on Stock Markets

TLDR: The research introduces VideoConviction, a multimodal dataset with expert annotations of finfluencer videos, designed to benchmark AI models’ ability to understand human conviction and stock recommendations. Findings show that while multimodal inputs improve stock ticker extraction, AI models still struggle with interpreting investment actions and conviction, highlighting a gap in financial reasoning compared to humans. A portfolio analysis reveals that following finfluencer recommendations generally underperforms market index funds, with an inverse strategy yielding higher returns but also higher risk.

In today’s digital age, financial influencers, or “finfluencers,” have become a significant presence on social media platforms like YouTube, sharing stock recommendations and investment advice. Understanding their true impact goes beyond just analyzing their words; it requires looking at multimodal signals such as their tone of voice, delivery style, and facial expressions. This is where a new research paper introduces a groundbreaking resource: VideoConviction.

The paper, titled “VideoConviction: A Multimodal Benchmark for Human Conviction and Stock Market Recommendations,” by Michael Galarnyk, Veer Kejriwal, Agam Shah, Yash Bhardwaj, Nicholas Watney Meyer, Anand Krishnan, and Sudheer Chava, addresses the need for a comprehensive dataset to evaluate how well artificial intelligence models can understand these complex human cues in financial discussions.

Introducing VideoConviction: A Unique Dataset

VideoConviction is a novel, expert-annotated multimodal dataset specifically designed for financial discourse. It comprises over 6,000 expert annotations, accumulated through 457 hours of human effort, from 288 YouTube videos featuring finfluencers discussing the U.S. stock market. Unlike previous datasets, VideoConviction is unique because it’s domain-specific to finance, includes expert annotations, and crucially, features a multimodal conviction score. This score, ranging from 1 to 3 (with 3 being the highest), assesses the finfluencer’s strength of belief based on their tone, facial expressions, overall delivery, and consistency with the video title.

The dataset includes both full-length video URLs with transcripts and segmented videos with transcripts, allowing for analysis at different levels of detail. This dual approach helps models understand recommendations within their immediate context and the broader video narrative. The creation of this dataset involved a meticulous process of curating finfluencer channels, filtering videos based on keywords, sampling to ensure temporal diversity, and then having expert annotators manually label stock recommendations, including ticker names, actions (buy, sell, hold, etc.), and the conviction score. Automatic speech recognition (ASR) was used to generate accurate transcripts for all videos and segments.

Benchmarking AI Models on Financial Understanding

The researchers used VideoConviction to benchmark various large language models (LLMs) and multimodal large language models (MLLMs). They evaluated these models on three sequential tasks: Task T (extracting the recommended stock ticker), Task TA (extracting the ticker and the associated action), and Task TAC (extracting the ticker, action, and conviction score). These tasks mirror the logical progression of understanding a financial recommendation, from identifying the stock to grasping the influencer’s conviction behind it.

The experiments revealed several key insights:

  • Multimodal inputs, which include video and audio alongside text, significantly improved the accuracy of ticker extraction (Task T). This is likely because visual cues, such as stock charts displayed in videos, often show the ticker names, helping models avoid errors.
  • However, models struggled considerably with more complex tasks like identifying investment actions (Task TA) and especially conviction (Task TAC). They often confused general market commentary with explicit buy/sell recommendations. Even with multimodal inputs, MLLMs found it challenging to fully capture the nuances of speaker intent, tone, and implicit reasoning, indicating a notable gap compared to human understanding in financial discourse.
  • Providing models with focused, segmented inputs generally led to better performance than using full-length videos. Full videos often contain irrelevant information like sponsorships or unrelated discussions, which can act as noise and hinder model effectiveness.
  • Despite the benefits of multimodal inputs for ticker extraction, their advantage diminished for tasks requiring deeper financial reasoning (TA and TAC). Interestingly, some LLMs, which only process text, performed comparably or even slightly better than MLLMs on the most complex task (TAC), suggesting that for certain reasoning-heavy financial tasks, LLMs might still be a more practical choice due to their lower cost and computational demands.

Analyzing Finfluencer Recommendations: A Portfolio Perspective

Beyond benchmarking AI models, the study also conducted a financial portfolio analysis using the expert-labeled data to understand the real-world implications of finfluencer recommendations. They evaluated several strategies a retail investor might use:

  • Buy-and-Hold: Investing in recommended stocks and holding them for six months.
  • Buy-and-Hold Weighted by Conviction: Allocating more capital to recommendations with higher conviction scores.
  • YouTuber Inverse: A contrarian approach, doing the opposite of the finfluencer’s advice (e.g., selling if they recommend buying).

These strategies were compared against popular market indices like the NASDAQ-100 (QQQ) and the S&P 500 (SPY).

The results were quite telling: simple buy-and-hold index strategies (QQQ and SPY) consistently outperformed finfluencer-based approaches. While the “Inverse YouTuber” strategy yielded the highest cumulative return, it also carried a higher risk. Furthermore, even high-conviction recommendations from finfluencers, while performing better than low-conviction ones, still lagged behind market benchmarks like QQQ. This suggests that relying solely on finfluencer advice, even when delivered with strong conviction, is often less effective than investing in broad market index funds. In fact, the analysis showed that 80% of finfluencer-recommended stocks failed to beat QQQ.

Also Read:

Conclusion and Future Directions

The VideoConviction benchmark provides a valuable tool for advancing research in multimodal financial analysis. It highlights that while multimodal AI models are crucial for tasks benefiting from visual and audio cues, text-based models can still be competitive for complex financial reasoning. The study also offers a crucial takeaway for investors: while finfluencers can be engaging, passive investment in market index funds often yields more reliable returns than following individual stock picks. The full research paper can be accessed here: VideoConviction Research Paper.

Future work in this area could focus on improving MLLMs’ ability to extract more empirical financial data, such as specific price targets, to further bridge the gap between AI understanding and human financial expertise.

Nikhil Patel
Nikhil Patelhttps://blogs.edgentiq.com
Nikhil Patel is a tech analyst and AI news reporter who brings a practitioner's perspective to every article. With prior experience working at an AI startup, he decodes the business mechanics behind product innovations, funding trends, and partnerships in the GenAI space. Nikhil's insights are sharp, forward-looking, and trusted by insiders and newcomers alike. You can reach him out at: [email protected]

- Advertisement -

spot_img

Gen AI News and Updates

spot_img

- Advertisement -