Topics / ETF and institutional holdings
ETF and institutional holdings
Holdings of the 500 largest US ETFs and five years of 13F filings for the 500 largest institutional managers.
Flow analysts, positioning research, fund-overlap studies.
16 listings on etf and institutional holdings
US 13F Institutional Holdings — Top 500 Funds (2021–2025)
Free
Five years of SEC Form 13F-HR institutional holdings for the top 500 US funds by AUM (combined AUM at Q4 2025 sample: $60.1T). 14,006,161 (manager × quarter × position) rows packaged as 20 per-quarter gzipped CSVs. Universe spans BlackRock, Vanguard / Geode, State Street, Fidelity (FMR), Berkshire Hathaway, Norges Bank, Bridgewater, Citadel, Renaissance, Two Sigma, Millennium, Coatue, Tiger Global, Pershing Square, Elliott, and ~485 more. Each row is one position held by one manager at quarter-end, with shares, USD value, filing date, and CUSIP. Read with pandas.readcsv(path, compression='gzip', parsedates=['periodenddate', 'filingdate']) or duckdb.readcsv('holdings_*.csv.gz'). Standard inputs for: hedge fund replication / cloning strategies, smart-money cluster signals, tracking famous-investor stake changes, sector rotation analysis, ownership-overlap network research, 13F-derived factor construction (popularity, conviction, concentration).
US ETF Holdings — Top 500 ETFs By AUM (Current Snapshot)
Free
Current holdings for the top 500 US-listed ETFs by AUM — 623,253 ETF-position rows covering 434 ETFs. Each row carries the parent ETF's symbol, name, AUM (market cap), sector and exchange, alongside the held asset's ticker, full name, ISIN, CUSIP, share count, weight percentage, and market value. Rows pre-sorted by ETF symbol then weight descending, so top holdings per fund appear first. Pairs naturally with the existing ETF Master List for fund-level metadata. Read with pandas.readcsv(path, compression='gzip', parsedates=['updatedAt']). Useful for: flow analysis (who's buying / dumping a stock at the ETF level), factor-exposure decomposition (roll up positions across multiple ETFs), copycat strategies for active ETFs, sector/factor overlap analysis, and identifying liquidity sources for specific names.
US Insider Trades — SEC Form 4 (Top ~1,000 Companies, 2021–2025)
Free
Five years of SEC Form 4 insider transaction filings for the top 918 US-listed companies by reporting history, covering 2021-01-04 through 2026-09-04. 601,965 transactions as a single gzipped CSV — one row per filed transaction. Columns include the filing date, transaction date, reporting insider name and role, transaction type (open-market buy/sell, option exercise, RSU vest, gift, etc.), shares transacted, price per share, total shares owned after the transaction, and a direct link to the EDGAR filing. Sourced from Financial Modeling Prep (Form 4 mirror of EDGAR). Read with pandas.readcsv(path, compression='gzip', parsedates=['filingDate', 'transactionDate']) or duckdb.read_csv('file.csv.gz'). Useful for: tracking insider sentiment, building cluster-buying signals, identifying executives unloading positions ahead of weakness, screen for management vs. board behavior divergence.
SEC Filings Index — Top 3,000 US Issuers, 10y (2016–2026)
Free
Ten years of SEC filing metadata for the top 3,000 US-listed companies by market cap. 2,167,270 filings (10-K, 10-Q, 8-K, S-1, DEF 14A, Form 4, 13F, etc.) as a single gzipped CSV. Each row has the symbol, CIK, filing date, accepted date, form type, and direct EDGAR URLs to both the index page and the final document. Read with pandas.readcsv(path, compression='gzip', parsedates=['filingDate', 'acceptedDate']). The 'where to find every regulatory filing' lookup table. Pair with the Earnings Transcripts and All US Quarterly Financials listings for full corporate-disclosure coverage. Use the link / finalLink URLs to fetch raw filing text from EDGAR for NLP / extraction pipelines (LLM training, 10-K MD&A topic modeling, 8-K material-event detection).
CFTC Commitment Of Traders — Weekly, 10y (2016–2026)
Free
Ten years of weekly Commitment of Traders (COT) reports from the CFTC for 65 futures contracts (E-Mini S&P, Nasdaq 100, gold, oil, natural gas, grains, metals, currency futures, Treasury futures, VIX, and more). 31,976 weekly rows in a single gzipped CSV with ~128 columns covering: long/short positions for commercial hedgers, non-commercial speculators (managed money), other reportables, and non-reportables (small specs/retail), plus net positioning, open interest, and percentage breakdowns. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). The classic positioning dataset for futures traders — useful for sentiment extremes, contrarian setups (commercial vs spec divergence), and macro-overlay strategies. Pair with the Macro Bundle's commodities OHLCV to combine price action with positioning shifts.
US Congressional Trading — STOCK Act Disclosures (Senate + House)
Free
Every Senate and House Periodic Transaction Report (PTR) filed under the STOCK Act, as indexed by FMP. 10,100 Senate trades + 10,100 House trades = 20,200 transactions in 2 gzipped CSVs. Each row has the politician's name, party district, transaction date, disclosure date (often weeks/months later — that lag is itself a signal), transaction type (Purchase / Sale / Exchange), asset symbol, asset description, asset type (Stock, Option, Bond, Mutual Fund), amount range (STOCK Act bins like '$1,001 – $15,000'), spouse/dependent owner flag, comment, and a direct URL to the official PDF/HTML disclosure. Read with pandas.readcsv(path, compression='gzip', parsedates=['transactionDate', 'disclosureDate']). Polymarket-trader catnip — track Pelosi, Crapo, Tuberville, et al. in near-real-time. Useful for: lawmaker-replication strategies, sector-rotation signals from committee members trading regulated industries, and detecting unusual transaction clusters around legislation.
Index Constituent Membership History — S&P 500, NASDAQ-100, Dow Jones
Free
Two-file companion dataset for survivorship-bias-free backtesting. (1) Current constituents (635 rows): every member of the S&P 500, NASDAQ-100, and Dow Jones Industrial Average as of build date, with sector, sub-sector, headquarters, founding year, CIK, and date first added to the index. (2) Historical changes (2055 rows): every add/drop event for these three indices, stretching back to 1957 (S&P 500), 1985 (NASDAQ-100), and 1994 (Dow Jones), with the symbol added, the symbol replaced, the date, and the reason given by S&P/NASDAQ/Dow. Why this matters: every realistic backtest of an index strategy on point-in-time membership requires knowing who was IN the index at each historical date — without this, you suffer survivorship bias (only seeing winners that survived to today). Reconstruct membership at any past date by starting with current and replaying changes backwards. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']).
SEC Company Tickers — Ticker To CIK And Registrant Name (All EDGAR Filers With A Ticker)
Free
10,412 ticker-to-CIK mappings for every SEC registrant with a listed ticker, straight from the SEC's company_tickers.json as of 2026-09-05. The join key between market data (tickers) and EDGAR filings (CIKs).
SEC XBRL Annual Fundamentals — 15 US-GAAP Concepts For Every Filer, Calendar Years 2009–2025
Free
1,152,906 filer-concept-year facts for 16,041 SEC registrants across 15 US-GAAP concepts (revenue, cost of revenue, gross profit, operating income, net income, total assets, liabilities, shareholders' equity, cash, long-term debt, operating cash flow, capital expenditure, diluted EPS and shares outstanding) for calendar years 2009 to 2025, from the SEC's XBRL frames API as of 2026-09-06. Long format: one row per company, concept and calendar-year frame with CIK, entity name, state or country of incorporation, unit, period start and end, value and the accession number of the filing the fact came from, 250 frames fetched. Join to the SEC Company Tickers set on CIK.
US ESG Scores — Disclosures, Ratings & Sector Benchmarks (Top 1,000)
Free
Environmental, Social, and Governance scores for the top 1,000 US-listed issuers by market cap (approximating the Russell 1000). Three complementary files: (1) esgdisclosures — 68,097 per-filing E/S/G/composite scores across 978 companies, linked back to the originating SEC form. (2) esgratings — 18,099 per-fiscal-year ESG risk-rating letters (A–F scale) plus industry rank, across 967 companies. (3) esgsectorbenchmark — 6,640 sector-level annual averages for benchmarking. ESG data of this quality is typically paywalled by MSCI / Sustainalytics; this dataset gives you an open, reproducible alternative. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). Useful for: ESG-tilted portfolio construction, sustainability research, regulatory disclosure analysis, and training models that need ESG features.
All US Quarterly Financials — IS / BS / CFS (2016–2025, ~10y)
Free
Ten years of quarterly Income Statement, Balance Sheet, and Cash Flow data for every USD-reporting US-listed company FMP indexes (~17,000 issuers including dual-class shares). 1,196,539 statement-period rows packaged as 3 gzipped CSV files, one CSV per statement type (income, balance, cashflow). Columns are the union of all FMP-reported fields per statement type — ~39 IS, ~61 BS, ~47 CFS. Read with pandas.readcsv(path, compression='gzip', parsedates=['reportPeriod', 'filingDate']) or duckdb.readcsv('incomestatement_*.csv.gz'). Coverage spans 2016 through 2025 quarterly. Useful for: fundamental factor research (gross margin, cash conversion, leverage), earnings drift studies, building screeners, fine-tuning LLMs on structured financial reporting, macro/sector-level financial-health rollups.
Macro Bundle — Treasury Yields, Forex, Commodities, Economic Indicators
Free
Four-file macro toolkit covering: (1) Treasury Yields: full daily yield curve since 1990 — 12 maturities (1mo through 30y) for 63 trading days. (2) Economic Indicators: 113 rows across 15 series (GDP, CPI, inflation rate, unemployment, federal funds, consumer sentiment, retail sales, industrial production, mortgage rates, jobless claims, nonfarm payrolls, durable goods, recession probabilities) since 1990 in long format. (3) Commodities OHLCV: 192,095 daily bars across 40 commodity contracts (E-Mini S&P, gold, oil, natural gas, grains, metals, etc.) 2006-present. (4) Forex OHLCV: 140,000 daily bars across 28 major + minor currency pairs (EURUSD, USDJPY, GBPUSD, AUDJPY, …) 2006-present. Read each with pandas.readcsv(path, compression='gzip', parsedates=['date']). Pair with the equity/crypto/fundamentals listings to build risk-on/risk-off regime models, macro-overlay strategies, currency-hedged backtests, or commodity-aware sector rotation.
Commodities Daily OHLCV — All FMP Contracts, 10 Years (2016–2026)
Free
Ten years of daily OHLCV across all 40 commodity contracts indexed by FMP — 109,218 rows in a single gzipped CSV. Covers energy (WTI/Brent crude, natural gas, gasoline, heating oil), precious & base metals (gold, silver, platinum, palladium, copper, aluminum), grains & oilseeds (wheat, corn, soybeans, rice, oats, soybean meal & oil), softs (sugar, coffee, cocoa, cotton, orange juice, lumber), livestock (live cattle, lean hogs, feeder cattle, class III milk), Treasury futures (2y, 5y, 10y, 30y), fed funds, US Dollar Index, and index futures (E-mini S&P, NASDAQ-100, mini Dow, Russell 2000). Each row tags a category field for fast filtering. Read with pandas.readcsv(path, compression='gzip', parsedates=['date']). Useful for: commodity-curve research, inflation-hedge backtests, cross-asset correlations, macro-regime detection, and rates/futures basis analysis.
US Analyst Activity — Ratings, Price Targets & Monthly Consensus (Top 3K, 5y)
Free
Five years of sell-side analyst activity for the top 3,000 US-listed issuers, delivered as three complementary files: 85,777 individual upgrade/downgrade events, 73,147 individual price target changes, and 144,410 monthly aggregate strong-buy / buy / hold / sell / strong-sell counts. Each event row carries the analyst firm, previous and new grade or target, the stock price at the time of the call, and a link back to the originating news item. Read with pandas.readcsv(path, compression='gzip', parsedates=['publishedDate']). Useful for: analyst-momentum factors, upgrade/downgrade event studies, calibrating analyst accuracy by firm, building target-change feature stacks, and reconstructing consensus drift around earnings.
FDIC-Insured Institutions — Every Active US Bank And Thrift With Assets, Deposits, Income And Regulator
Free
4,235 active FDIC-insured institutions with certificate number, location, bank class, primary regulator, Fed RSSD id, establishment date, office count, total assets, deposits, net income, ROA and ROE, from the FDIC BankFind API as of 2026-09-05.
FRED Macro Core — 36 US Rates, Inflation, Labour, Activity And Fed Series (Daily To Quarterly)
Free
200,021 observations across 36 core US macro series from FRED: the Treasury curve (1-month to 30-year) and curve spreads, fed funds and SOFR, TIPS and breakeven inflation, high-yield and investment-grade OAS, the broad dollar index and major FX crosses, WTI and Henry Hub, CPI, core CPI and PCE, unemployment, payrolls, initial claims, industrial production, retail sales, housing starts, consumer sentiment, nominal and real GDP, M2, the Fed balance sheet, vehicle sales and real disposable income. Long format (series_id, date, value) with series name, frequency and unit on every row, full history to 2026-09-05.
Get this data into your agent
Point any MCP client (Claude, Cursor, your own agent) at https://dagentbase.com/api/mcp with an API key as the bearer token and it can search, preview and claim every listing above in one call. The REST API, the TypeScript and Python SDKs and per-listing markdown pages (/listing/{id}.md) cover everything else.