Paper: Leveraging Large Language Models to Bridge On-chain and Off-chain Transparency in Stablecoins
Authors: Yuexin Xiang, Yuchen Lei, SM Mahir Shazeed Rish, Yuanzhe Zhang, Qin Wang, Tsz Hon Yuen, Jiangshan Yu
Date: 2 Dec 2025
Estimated Reading Time: 20 minutes
This paper addresses the transparency gap in stablecoins arising from the separation between verifiable on-chain data and unstructured off-chain disclosures. It proposes an automated framework using large language models to semantically align issuer attestations with blockchain-observed issuance and market data. The framework integrates multi-chain on-chain metrics and issuer disclosures through a Model Context Protocol that standardizes access and temporal alignment. An LLM-based analytical layer extracts reserve and liability indicators from disclosures and compares them against observed circulation and price behavior. Empirical analysis focuses on USDT and USDC from 2022 to 2024, identifying discrepancies between reported reserves and on-chain data. The results show systematic differences in disclosure quality and timing across issuers and demonstrate how LLM-assisted analysis can support automated auditing. Overall, the paper positions LLMs as a tool for scalable, cross-modal transparency assessment in stablecoins.
Core insights
- Cross-modal transparency gap: Stablecoin transparency is fragmented between on-chain issuance data and off-chain textual disclosures that are not natively connected. This separation limits independent verification and creates informational asymmetries.
- LLM-based semantic alignment: Large language models can parse unstructured attestations and map extracted financial indicators to corresponding on-chain metrics. This enables systematic comparison of reported reserves with observed circulation.
- Model Context Protocol (MCP): The MCP standardizes access to heterogeneous data sources and aligns them on a common temporal axis. This design supports reproducible and auditable analysis across quantitative and qualitative inputs.
- Issuer heterogeneity: Empirical results show that USDT and USDC maintain similar peg stability but differ materially in disclosure cadence, reserve reporting regularity, and transparency consistency. These differences affect how market stress is reflected in disclosures.
- Automated auditing potential: LLM-assisted analysis can classify disclosure periods as normal, suspicious, or abnormal based on reserve coverage and market consistency. This supports scalable monitoring of stablecoin integrity.
The framework treats stablecoin token supply as observable on-chain while reserve backing is disclosed off-chain, making transparency a function of alignment rather than absolute values. By extracting reported circulation, assets, and liabilities from attestations, the system compares them against observed market capitalization and supply. This approach emphasizes consistency between reported backing and circulating supply as a core determinant of credibility.
From a tokenomics perspective, stablecoin demand depends on confidence in redeemability, which is influenced by both liquidity and disclosure quality. The analysis shows that USDT maintains high demand through scale and liquidity even when disclosures are irregular, while USDC relies more heavily on predictable attestations to sustain confidence. This raises the question of whether market demand is more sensitive to liquidity depth or to disclosure precision during stress periods.
The MCP architecture shifts transparency from a narrative exercise to a structured data problem. By enforcing typed interfaces and deterministic retrieval, it reduces ambiguity in how disclosures are interpreted relative to on-chain data. However, the framework assumes that issuer disclosures, once parsed, are truthful representations of reserves. How robust is the system if disclosures are systematically incomplete or strategically vague?
The empirical classification of abnormal and suspicious periods illustrates how supply gaps and turnover spikes correlate with disclosure timing. In May 2022 and March 2023, deviations were linked to delayed or stress-related disclosures rather than immediate reserve shortfalls. This suggests that timing mismatches, rather than absolute reserve insufficiency, can drive temporary loss of confidence. A further question is whether shorter disclosure intervals would materially reduce volatility in such episodes.
The reliance on rule-based thresholds for classification ensures interpretability but may overlook nonlinear dynamics between issuance, redemption, and price. For example, large but brief supply gaps may have different implications than smaller but persistent ones. Extending the framework to incorporate adaptive thresholds could change how transparency risk is assessed over time.
Overall, the paper frames transparency as an endogenous component of stablecoin stability rather than an external reporting feature. By operationalizing disclosure alignment through LLMs, it shows how automated analysis can support ongoing monitoring of supply backing and market confidence. The results suggest that scalable transparency tools can indirectly support stablecoin demand by reducing uncertainty around reserve adequacy and disclosure practices.
