Topics / SEC filings and EDGAR
SEC filings and EDGAR
Ten years of the EDGAR filings index for the 3,000 largest US issuers, the ticker-to-CIK map for every registrant, and XBRL fundamentals straight from the SEC's frames API.
Compliance and legal teams, NLP pipelines over filings, fundamental researchers.
17 listings on sec filings and edgar
SEC Company Tickers — Ticker To CIK And Registrant Name (All EDGAR Filers With A Ticker)
Free
10,412 ticker-to-CIK mappings for every SEC registrant with a listed ticker, straight from the SEC's company_tickers.json as of 2026-09-05. The join key between market data (tickers) and EDGAR filings (CIKs).
SEC Filings Index — Top 3,000 US Issuers, 10y (2016–2026)
Free
Ten years of SEC filing metadata for the top 3,000 US-listed companies by market cap. 2,167,270 filings (10-K, 10-Q, 8-K, S-1, DEF 14A, Form 4, 13F, etc.) as a single gzipped CSV. Each row has the symbol, CIK, filing date, accepted date, form type, and direct EDGAR URLs to both the index page and the final document. Read with pandas.readcsv(path, compression='gzip', parsedates=['filingDate', 'acceptedDate']). The 'where to find every regulatory filing' lookup table. Pair with the Earnings Transcripts and All US Quarterly Financials listings for full corporate-disclosure coverage. Use the link / finalLink URLs to fetch raw filing text from EDGAR for NLP / extraction pipelines (LLM training, 10-K MD&A topic modeling, 8-K material-event detection).
SEC XBRL Annual Fundamentals — 15 US-GAAP Concepts For Every Filer, Calendar Years 2009–2025
Free
1,152,906 filer-concept-year facts for 16,041 SEC registrants across 15 US-GAAP concepts (revenue, cost of revenue, gross profit, operating income, net income, total assets, liabilities, shareholders' equity, cash, long-term debt, operating cash flow, capital expenditure, diluted EPS and shares outstanding) for calendar years 2009 to 2025, from the SEC's XBRL frames API as of 2026-09-06. Long format: one row per company, concept and calendar-year frame with CIK, entity name, state or country of incorporation, unit, period start and end, value and the accession number of the filing the fact came from, 250 frames fetched. Join to the SEC Company Tickers set on CIK.
US Insider Trades — SEC Form 4 (Top ~1,000 Companies, 2021–2025)
Free
Five years of SEC Form 4 insider transaction filings for the top 918 US-listed companies by reporting history, covering 2021-01-04 through 2026-09-04. 601,965 transactions as a single gzipped CSV — one row per filed transaction. Columns include the filing date, transaction date, reporting insider name and role, transaction type (open-market buy/sell, option exercise, RSU vest, gift, etc.), shares transacted, price per share, total shares owned after the transaction, and a direct link to the EDGAR filing. Sourced from Financial Modeling Prep (Form 4 mirror of EDGAR). Read with pandas.readcsv(path, compression='gzip', parsedates=['filingDate', 'transactionDate']) or duckdb.read_csv('file.csv.gz'). Useful for: tracking insider sentiment, building cluster-buying signals, identifying executives unloading positions ahead of weakness, screen for management vs. board behavior divergence.
US ESG Scores — Disclosures, Ratings & Sector Benchmarks (Top 1,000)
Free
Environmental, Social, and Governance scores for the top 1,000 US-listed issuers by market cap (approximating the Russell 1000). Three complementary files: (1) esgdisclosures — 68,097 per-filing E/S/G/composite scores across 978 companies, linked back to the originating SEC form. (2) esgratings — 18,099 per-fiscal-year ESG risk-rating letters (A–F scale) plus industry rank, across 967 companies. (3) esgsectorbenchmark — 6,640 sector-level annual averages for benchmarking. ESG data of this quality is typically paywalled by MSCI / Sustainalytics; this dataset gives you an open, reproducible alternative. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). Useful for: ESG-tilted portfolio construction, sustainability research, regulatory disclosure analysis, and training models that need ESG features.
S&P 500 Daily OHLCV (2006–2026, 20 Years)
Free
Twenty years of daily open/high/low/close/volume bars for every current S&P 500 constituent (503 tickers including dual-class shares like BRK.A/BRK.B and GOOG/GOOGL). Date range: 2006-01-03 through 2026-09-05. Sourced live from Financial Modeling Prep and packaged as a single gzipped CSV — one row per (symbol, date), sorted by symbol then date ascending. Pandas, DuckDB, R, and most data tools read .csv.gz natively. Bedrock dataset for backtesting, factor research, event studies, and ML training. Note: prices reflect FMP's reported values at fetch time and are dividend/split adjusted as provided by FMP.
US Congressional Trading — STOCK Act Disclosures (Senate + House)
Free
Every Senate and House Periodic Transaction Report (PTR) filed under the STOCK Act, as indexed by FMP. 10,100 Senate trades + 10,100 House trades = 20,200 transactions in 2 gzipped CSVs. Each row has the politician's name, party district, transaction date, disclosure date (often weeks/months later — that lag is itself a signal), transaction type (Purchase / Sale / Exchange), asset symbol, asset description, asset type (Stock, Option, Bond, Mutual Fund), amount range (STOCK Act bins like '$1,001 – $15,000'), spouse/dependent owner flag, comment, and a direct URL to the official PDF/HTML disclosure. Read with pandas.readcsv(path, compression='gzip', parsedates=['transactionDate', 'disclosureDate']). Polymarket-trader catnip — track Pelosi, Crapo, Tuberville, et al. in near-real-time. Useful for: lawmaker-replication strategies, sector-rotation signals from committee members trading regulated industries, and detecting unusual transaction clusters around legislation.
Index Constituent Membership History — S&P 500, NASDAQ-100, Dow Jones
Free
Two-file companion dataset for survivorship-bias-free backtesting. (1) Current constituents (635 rows): every member of the S&P 500, NASDAQ-100, and Dow Jones Industrial Average as of build date, with sector, sub-sector, headquarters, founding year, CIK, and date first added to the index. (2) Historical changes (2055 rows): every add/drop event for these three indices, stretching back to 1957 (S&P 500), 1985 (NASDAQ-100), and 1994 (Dow Jones), with the symbol added, the symbol replaced, the date, and the reason given by S&P/NASDAQ/Dow. Why this matters: every realistic backtest of an index strategy on point-in-time membership requires knowing who was IN the index at each historical date — without this, you suffer survivorship bias (only seeing winners that survived to today). Reconstruct membership at any past date by starting with current and replaying changes backwards. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']).
US Stock News Archive — Top 1,500 Companies, 5y (2021–2026)
Free
Five years of ticker-tagged news articles for the top 1,500 US-listed issuers by market cap, covering 1,453,949 articles across 1,453 companies. Each row carries publication timestamp, publisher, site domain, headline, article body snippet, canonical URL, and image URL — all linked back to the underlying ticker symbol. Split into one CSV.GZ file per calendar quarter (21 files) for partial-history loading. Read with pandas.readcsv(path, compression='gzip', parsedates=['publishedDate']). Useful for: training news-sentiment models, event-study backtests around news catalysts, building ticker-mention frequency time series, fine-tuning LLMs on financial-news prose, and reconstructing the news flow around earnings, M&A, regulatory events, and macro shocks.
US Key Ratios & Metrics — Quarterly, Top 3,000, 10y (2016–2026)
Free
Ten years of pre-computed quarterly financial ratios and metrics for the top 3,000 US-listed issuers, covering 103,013 ratio-period rows across 2,973 companies and 102,983 key-metric rows across 2,973 companies. Two complementary files in one listing: (1) keyratiosquarterly — profitability margins, turnover ratios, liquidity, solvency, leverage, valuation multiples (P/E, P/B, P/S, P/FCF, EV/EBITDA), per-share book values, dividend metrics, and tax/interest burdens. (2) keymetricsquarterly — market cap, enterprise value, EV multiples, return-on-capital family (ROA, ROE, ROIC, ROCE), working-capital cycles (DSO, DPO, DIO, cash conversion cycle), and capex/R&D/SBC intensity. Saves buyers weeks of feature engineering on top of raw IS/BS/CFS statements. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). Useful for: fundamental factor models, multi-factor backtests, screening by financial health, training ML models on pre-computed financial features.
US 13F Institutional Holdings — Top 500 Funds (2021–2025)
Free
Five years of SEC Form 13F-HR institutional holdings for the top 500 US funds by AUM (combined AUM at Q4 2025 sample: $60.1T). 14,006,161 (manager × quarter × position) rows packaged as 20 per-quarter gzipped CSVs. Universe spans BlackRock, Vanguard / Geode, State Street, Fidelity (FMR), Berkshire Hathaway, Norges Bank, Bridgewater, Citadel, Renaissance, Two Sigma, Millennium, Coatue, Tiger Global, Pershing Square, Elliott, and ~485 more. Each row is one position held by one manager at quarter-end, with shares, USD value, filing date, and CUSIP. Read with pandas.readcsv(path, compression='gzip', parsedates=['periodenddate', 'filingdate']) or duckdb.readcsv('holdings_*.csv.gz'). Standard inputs for: hedge fund replication / cloning strategies, smart-money cluster signals, tracking famous-investor stake changes, sector rotation analysis, ownership-overlap network research, 13F-derived factor construction (popularity, conviction, concentration).
Global Index Daily OHLCV — 20 Years (2006–2026)
Free
Twenty years of daily OHLCV for 36 major global equity indices, volatility benchmarks, and US Treasury-yield series — 185,681 index-day rows in a single gzipped CSV. Coverage spans US broad-market (S&P 500/400/600, NASDAQ Composite, NASDAQ-100, Dow, Russell 2000, NYSE), volatility (VIX, VXN), Treasury yields (5y/10y/30y), European majors (FTSE 100, DAX, CAC 40, Euro Stoxx 50, IBEX, FTSE MIB, SMI, AEX, OMX), Asia-Pacific (Nikkei 225, TOPIX, Hang Seng, Shanghai/Shenzhen, KOSPI, TAIEX, SENSEX, NIFTY 50, ASX 200, NZX 50), and Americas ex-US (TSX, Bovespa, IPC, Merval). Pairs naturally with the existing Index Constituent History listing for survivorship-bias-free benchmark studies. Each row tags symbol, human-readable name, country (ISO), and quote currency. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). Useful for: benchmark backtesting, beta computation, regime detection (VIX/yields), cross-asset correlation studies, and as the reference series for relative-strength signals.
US Earnings Call Transcripts (Top 1,000 Companies, 2021–2025)
Free
Five years of full-text earnings call transcripts from the top ~1,000 US public companies by reporting frequency (plus Tesla and Meta, which IPO'd more recently), covering Q1 2021 through Q4 2025. 19,677 transcripts across 1002 symbols, packaged as 10 gzipped JSONL files — one file per half-year (H1 = Q1+Q2, H2 = Q3+Q4) to stay under platform per-file size limits. Each line is one transcript with fields: symbol, fiscalyear, fiscalquarter, date, content (full prepared remarks + Q&A, speaker labels preserved). Sourced from Financial Modeling Prep. Read natively with pandas.readjson(path, lines=True, compression='gzip') or duckdb.readjson('file.jsonl.gz', lines=true) — concatenate all files for the full 5-year corpus. Industry-grade NLP corpus: ~25 KB per transcript average. Ideal for sentiment models, topic modeling, executive language change-detection, earnings-drift research keyed off transcript embedding similarity, and LLM fine-tuning on real corporate disclosure language.
All US Quarterly Financials — IS / BS / CFS (2016–2025, ~10y)
Free
Ten years of quarterly Income Statement, Balance Sheet, and Cash Flow data for every USD-reporting US-listed company FMP indexes (~17,000 issuers including dual-class shares). 1,196,539 statement-period rows packaged as 3 gzipped CSV files, one CSV per statement type (income, balance, cashflow). Columns are the union of all FMP-reported fields per statement type — ~39 IS, ~61 BS, ~47 CFS. Read with pandas.readcsv(path, compression='gzip', parsedates=['reportPeriod', 'filingDate']) or duckdb.readcsv('incomestatement_*.csv.gz'). Coverage spans 2016 through 2025 quarterly. Useful for: fundamental factor research (gross margin, cash conversion, leverage), earnings drift studies, building screeners, fine-tuning LLMs on structured financial reporting, macro/sector-level financial-health rollups.
Crypto Daily OHLCV — Top 250 By Market Cap (2020–2026, 5y)
Free
Five years of daily OHLCV bars for the 250 largest cryptocurrencies by market cap as of build date — Bitcoin, Ethereum, Tether, BNB, XRP, Solana, USDC and 243 others. 426,989 daily bars as a single gzipped CSV — one row per (symbol, date), sorted by symbol then date ascending. Date range covers the 2020 bull, 2022 bear, 2024 BTC ETF approval, 2025 cycle peak. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']) or duckdb.readcsv('cryptoohlcv_5y.csv.gz'). Bedrock dataset for: crypto strategy backtesting, regime detection, BTC dominance studies, alt-season identification, vol-of-vol analysis, training ML models on cycles 2020-2025.
All US Stock Daily OHLCV — Top 3,000 By Market Cap (2016–2026, 10y)
Free
Ten years of daily open/high/low/close/volume bars for the top 3,000 US-listed stocks by market cap (NVIDIA, GOOGL/GOOG, AAPL, MSFT, … through ~$1B mid-caps). 6,148,886 rows packaged as 11 per-year gzipped CSVs. Universe spans NASDAQ, NYSE, and AMEX listings only (filtered out foreign exchanges and ADRs with dot-suffixes). Read with pandas.readcsv(path, compression='gzip', parsedates=['date']) or duckdb.readcsv('usstocksohlcv*.csv.gz') to load all years at once. Pair this with the Index Constituent Membership History listing for survivorship-bias-free Russell-3000-style backtesting. Bedrock dataset for full-universe quant research, factor model construction, and ML training across small/mid/large-cap regimes.
US Analyst Activity — Ratings, Price Targets & Monthly Consensus (Top 3K, 5y)
Free
Five years of sell-side analyst activity for the top 3,000 US-listed issuers, delivered as three complementary files: 85,777 individual upgrade/downgrade events, 73,147 individual price target changes, and 144,410 monthly aggregate strong-buy / buy / hold / sell / strong-sell counts. Each event row carries the analyst firm, previous and new grade or target, the stock price at the time of the call, and a link back to the originating news item. Read with pandas.readcsv(path, compression='gzip', parsedates=['publishedDate']). Useful for: analyst-momentum factors, upgrade/downgrade event studies, calibrating analyst accuracy by firm, building target-change feature stacks, and reconstructing consensus drift around earnings.
Get this data into your agent
Point any MCP client (Claude, Cursor, your own agent) at https://dagentbase.com/api/mcp with an API key as the bearer token and it can search, preview and claim every listing above in one call. The REST API, the TypeScript and Python SDKs and per-listing markdown pages (/listing/{id}.md) cover everything else.