50
Listings
0
Sales
—
Avg Rating
93/100
Avg Quality
Free
Two-file dataset of every dividend payment and every stock split for every USD-reporting US-listed issuer FMP indexes (~17,500). 410,299 dividend rows + 17,982 split rows. Dividend rows include declaration / record / payment dates, the actual dividend amount, the split-adjusted dividend, the yield at announcement, and frequency (Quarterly, Monthly, Annual, Special). Split rows include the date and the split ratio (numerator / denominator). Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). Required for: building total-return series from price-only OHLCV (combine with the S&P 500 OHLCV dataset), dividend-growth screening, post-split adjustment of historical prices, special-dividend event studies.
Free
Five years of SEC Form 13F-HR institutional holdings for the top 500 US funds by AUM (combined AUM at Q4 2025 sample: $60.1T). 14,006,161 (manager × quarter × position) rows packaged as 20 per-quarter gzipped CSVs. Universe spans BlackRock, Vanguard / Geode, State Street, Fidelity (FMR), Berkshire Hathaway, Norges Bank, Bridgewater, Citadel, Renaissance, Two Sigma, Millennium, Coatue, Tiger Global, Pershing Square, Elliott, and ~485 more. Each row is one position held by one manager at quarter-end, with shares, USD value, filing date, and CUSIP. Read with pandas.readcsv(path, compression='gzip', parsedates=['periodenddate', 'filingdate']) or duckdb.readcsv('holdings_*.csv.gz'). Standard inputs for: hedge fund replication / cloning strategies, smart-money cluster signals, tracking famous-investor stake changes, sector rotation analysis, ownership-overlap network research, 13F-derived factor construction (popularity, conviction, concentration).
Free
Five years of ticker-tagged news articles for the top 1,500 US-listed issuers by market cap, covering 1,453,949 articles across 1,453 companies. Each row carries publication timestamp, publisher, site domain, headline, article body snippet, canonical URL, and image URL — all linked back to the underlying ticker symbol. Split into one CSV.GZ file per calendar quarter (21 files) for partial-history loading. Read with pandas.readcsv(path, compression='gzip', parsedates=['publishedDate']). Useful for: training news-sentiment models, event-study backtests around news catalysts, building ticker-mention frequency time series, fine-tuning LLMs on financial-news prose, and reconstructing the news flow around earnings, M&A, regulatory events, and macro shocks.