TLDR: A new study investigates whether large language models (LLMs) can effectively make “moral” judgments in real-world, high-stakes scenarios like allocating homelessness resources. The research found that LLM judgments were highly inconsistent internally, did not align with established vulnerability scoring systems, and were poor predictors of actual human caseworker decisions. The findings suggest that current LLMs are not ready for direct integration into critical societal decision-making without significant domain-specific adaptation and human oversight.
Large Language Models (LLMs) are increasingly being explored for their ability to make complex judgments, including those with ethical and societal implications. While much of the discussion has focused on their alignment with human judgments in theoretical scenarios, a recent study delves into a more immediate and critical application: their readiness to assist or even replace ‘street-level bureaucrats’ – individuals who make decisions about allocating scarce social resources, such as housing for people experiencing homelessness.
The research, titled “Street-Level AI: Are Large Language Models Ready for Real-World Judgments?” by Gaurab Pokharel, Shafkat Farabi, Patrick J. Fowler, and Sanmay Das, highlights that these real-world scenarios carry much higher stakes than hypothetical dilemmas. Decisions in areas like homelessness services, post-disaster medical triage, or organ transplantation directly impact people’s lives and require nuanced understanding of local justice principles and established prioritization practices.
The Challenge of Real-World Allocation
Currently, resource allocation in social services often relies on a combination of standardized scoring systems and the discretion of experienced caseworkers. For instance, homelessness services frequently use a “vulnerability first” approach, prioritizing those at greatest risk based on tools like the Vulnerability Index-Service Prioritization Decision Assistance Tool (VI-SPDAT). These systems, however, are complemented by the on-the-ground knowledge and judgment of street-level bureaucrats, who navigate complex individual needs and contextual nuances.
The study aimed to understand how LLM judgments align with both human judgments and these established bureaucratic scoring systems. Crucially, it used real data from individuals needing services, ensuring strict confidentiality by using local, large models for analysis.
How LLMs Were Tested
The researchers conducted two main types of experiments:
1. Pairwise Comparisons: LLMs were presented with pairs of household profiles (including demographics, income, disability, and service requests) and asked to prioritize one for more intensive transitional housing, similar to how human subjects were tested in previous work. This helped assess if LLMs leaned towards prioritizing based on vulnerability or potential positive outcomes.
2. Ranking Task: Using real-world vulnerability assessment data from St. Louis, LLMs were tasked with creating a complete ranked list of households for prioritization. These LLM-generated rankings were then compared against established bureaucratic scoring systems like the VI-SPDAT and the Risk/Medical Frailty Score (RMFS).
Key Findings: Inconsistency and Misalignment
The results raised significant concerns about the immediate deployment of off-the-shelf LLMs in such high-stakes domains:
- Internal Inconsistency: In the ranking task, LLMs (specifically LL-3-8B and DS-7B) showed low consistency across independent runs. This means that the same model, given the exact same data, could produce significantly different prioritization rankings, undermining reliability.
- Lack of Alignment with Bureaucratic Systems: LLM-generated rankings had near-zero, and sometimes even negative, correlation with the established VI-SPDAT and RMFS scores. This indicates that LLMs do not capture the vulnerability principles embedded in existing, socially and politically determined systems.
- Poor Prediction of Human Decisions: When comparing LLM rankings to actual caseworker decisions on service allocation, the LLMs were found to be weak predictors. They offered no improvement over existing bureaucratic tools in forecasting real-world prioritization decisions made by human experts.
- Inconsistent Feature Focus: An analysis of which questionnaire responses most influenced LLM decisions revealed that models focused on different features across runs, even showing inconsistent polarity (whether a feature was favorable or adverse). This suggests a struggle to identify consistent criteria for judging vulnerability.
While LLMs in pairwise comparisons showed some qualitative similarities to non-expert human judgments (e.g., variability in orientation without explicit risk information), their performance in the more complex ranking task with real-world data highlighted serious limitations.
Also Read:
- Why Large Language Models Can’t Replace Human Participants in Psychological Research
- AI Models Master Community Resource Allocation Through Participatory Budgeting
Implications for AI in Social Services
The study concludes that current generation AI systems are not ready for naive integration into high-stakes societal decision-making. The pronounced inconsistency, disconnect from established vulnerability metrics, and failure to capture the nuanced discretion of experienced caseworkers pose significant risks. Relying on LLMs without domain-specific adaptation could lead to inefficient resource deployment, erode community-driven prioritization principles, and potentially exacerbate service gaps and inequities.
The researchers emphasize the need for rigorous, context-grounded evaluation before automating public resource allocation. Future work should explore strategies like fine-tuning LLMs on localized caseworker data, integrating multi-modal client information, or embedding human-in-the-loop safeguards. Ultimately, any AI augmentation in this domain must reflect the complex trade-offs and moral frameworks that have evolved through community engagement and political processes, rather than unmediated reliance on present-day language models. For more details, you can read the full paper here.


