The best AI projects for a finance student check a machine's work on real, public finance data. Ask an AI tool 40 questions about a company's annual report and grade every answer against the filing. Let one build a valuation, then audit it cell by cell. Score a sentiment model against human labels, judge a fraud model by the frauds it catches, and rebuild a published recession model before claiming yours is better.

What makes each one worth showing is not the model call. Anyone with a login can make that. It is the table at the end: where the model was wrong, and why.

An AI project, for a finance student, is a small finished piece of work that puts an AI model on real financial data and shows, with numbers, where the model was right and where it failed.

The five projects at a glance

Each runs on data a student may use, and each teaches a check the other four do not.

ProjectDataToolsWhat it provesClosest seat
Audit an AI's answers against a 10-KSEC EDGAR filings and XBRL data; the FinanceBench sampleA chat model, a spreadsheet or PythonYou know where each number in a filing livesBanking, equity research, credit
Rebuild an AI-generated valuationThe company's 10-K; the 10-year Treasury yieldExcel and an AI assistantYou review a model the way a senior reviews a junior'sBanking, private equity
Test a sentiment model against human labelsFinancial PhraseBank; the Loughran-McDonald word listsA FinBERT model and a chat modelYou can evaluate language models on finance textResearch, markets, asset management
Judge a fraud model on what it catchesThe ULB credit card fraud datasetPython and scikit-learnYou measure models honestly on rare eventsRisk, credit, bank data science
Replicate a published recession modelSt. Louis Fed economic series and their archived vintages; NBER datesExcel, then PythonYou match a published model before claiming an edgeMarkets, macro, asset management

What makes an AI project worth building?

Almost anyone can now get a model to produce a summary, a forecast or a comps table, which is why producing one proves so little. OpenAI's own help page on whether ChatGPT tells the truth puts the catch in one line: "Use ChatGPT as a first draft, not a final source."

Two banks that built their own tools describe the same split. JPMorgan's chief analytics officer talks about junior staff moving from makers to checkers, and Morgan Stanley ran evals on its wealth management tools, grading the machine's output against expert answers before a tool shipped. A project that shows you doing the grading is practice for the job as it now exists.

The demo

  • A chatbot that answers questions about any stock
  • A price predictor with no benchmark
  • A summary tool nobody checked against the source
  • A dashboard of model outputs

The same project, checked

  • The chatbot, graded on 40 questions with known answers
  • The predictor, scored out of sample against a simple benchmark
  • The summaries, each line traced back to the filing
  • The dashboard, with an error rate and a log of what went wrong

None of the demos is worthless: each crosses over the moment it carries a test set, a benchmark and an honest error rate. The five projects below are built that way from the start, and any one of them is the small checking project that closes the preparation list for the changing analyst job.

Project 1: Audit an AI's answers against a 10-K

The first project is the purest version of the idea: ask a model questions whose answers sit in a public filing, and grade every answer. It teaches where numbers live, because a revenue figure is only right for one period, under one tag, in one unit.

Where the answers live: EDGAR and FinanceBench

US public companies' filings sit on the SEC's EDGAR system, and the SEC's access page is blunt about who may use them: "Anyone can access and download this information for free." The structured numbers are there too. Its XBRL APIs return the financial statement data from 10-Ks, 10-Qs and other forms as JSON, and "do not require any authentication or API keys to access."

For a ready-made answer key, FinanceBench, a November 2023 test set built from 10-Ks, 10-Qs, 8-Ks and earnings reports, publishes an open sample of 150 questions with human-written answers and the page each comes from, and how leading models scored on it is a story of its own.

Forty questions, asked twice

Pick one company and write 40 questions a first-year analyst would be asked: revenue, segment profit, share count, debt maturing next year, the change in operating margin. Record each right answer and where it sits. Then ask a model every question twice, once from memory and once with the filing attached, and log each answer.

A spreadsheet works for the first version. A hosted notebook such as Google Colab, which "requires no setup to use," is enough to pull the XBRL numbers with a few lines of code.

Four traps in Apple's XBRL data

Every example below comes from Apple's own XBRL data at the SEC.

CheckWhat goes wrongApple example
PeriodFiscal and calendar years get mixedFiscal 2025 ran from September 29, 2024 to September 27, 2025
TagA script queries the wrong XBRL conceptThe generic Revenues tag stops at fiscal 2018, at $265.6bn
FrameA calendar label hides a fiscal yearThe SEC's calendar frames file fiscal 2024, which ended September 28, 2024, as CY2024
UnitsWhole dollars read as millions, or the reverseXBRL stores fiscal 2025 revenue as 416,161,000,000

Put the first check to work. Apple's revenue was $416.2bn in fiscal 2025 against $391.0bn in fiscal 2024: 416.2 divided by 391.0 is 1.064, so growth was 6.4%. Ask a model for "Apple's 2025 revenue" and a good answer names the fiscal year. A weak one gives a number and leaves you to guess which twelve months it covers.

The tag check matters for code rather than chat. Apple's revenue sits under RevenueFromContractWithCustomerExcludingAssessedTax. A script that asks the SEC for Apple's Revenues instead, the kind an AI assistant writes in seconds, gets back a figure from fiscal 2018, with no error message, because the query ran perfectly.

The finished audit holds the 40 questions, the model's answers and the correct values with their page or tag, an error log sorted by the four traps, and the three most instructive misses. Its headline is the error rate with the filing attached and without it. Engineers can go a step further and grade a retrieval pipeline on the same questions, the build-then-grade exercise that prepares an engineer for AI roles at hedge funds.

Test yourself

Interview level

An AI-written script asks the SEC's XBRL data for Apple's Revenues concept. Why does it return a fiscal 2018 figure?

Project 2: Rebuild an AI-generated valuation model

The second project moves the habit from a filing to a spreadsheet. A filing has one right revenue figure; a valuation has no single right answer, only assumptions someone has to defend, so the check falls on the model's structure and on the inputs nobody sourced.

Inputs: the 10-K, FRED and Damodaran

The company's 10-K supplies the history. For the risk-free rate, take the 10-year Treasury constant-maturity yield, series DGS10 on FRED, the St. Louis Fed's database of hundreds of thousands of economic series.

For an equity risk premium, Aswath Damodaran, who teaches corporate finance and valuation at New York University's Stern School of Business, publishes implied premiums for the US market and refreshes his data in the first two weeks of each year.

The reviewer's checklist

Ask an AI assistant, in Excel or in a chat window, to build a five-year discounted cash flow model of one company from its 10-K; Copilot in Excel is one option, and any assistant works. Then stop building and start reviewing:

  1. Trace every hard-coded number to a page of the 10-K, or flag it as the tool's own assumption.
  2. Check that the balance sheet balances in every forecast year, and that cash on the cash flow statement matches cash on the balance sheet.
  3. Confirm units and signs. The ICAEW's Financial Modelling Code, the accountancy body's 2024 set of principles, asks for clear units and "a clear sign convention."
  4. Read the terminal value last, because one assumption there can swing the answer.

Worked check: the terminal value

Take a terminal value built with the standard growth formula: the final forecast year's cash flow, grown by one more year, divided by the discount rate minus growth. At an 8% discount rate and 3% growth, that is 1.03 divided by 0.05, or 20.6 times the final year's cash flow. Let the tool pick 6% instead and it becomes 1.06 divided by 0.02, or 53 times: about 2.6 times larger, from one unsourced number.

The rates are illustrative; the arithmetic is not. The check is a question rather than a formula: where did 6% come from, and can any company outgrow the whole economy forever? Asking it is the reviewer's job.

The finished project is two files and a log: the tool's original model, your corrected version, and an audit log listing each issue, its cell, the fix and its effect on value per share.

Test yourself

Interview level

At an 8% discount rate, what happens to a growth-formula terminal value when assumed growth rises from 3% to 6%?

Project 3: Test an AI sentiment model against human labels

Filings have right answers, and a model's structure can be tested. Language is harder. The third project asks what to do when the answer key itself is open to argument, by scoring machines that read financial news against people who did the same job.

Financial PhraseBank and where its labellers disagree

Financial PhraseBank holds about 4,800 sentences from English-language financial news, labelled positive, negative or neutral by a pool of 16 people with a background in financial markets, 13 of them master's students at Aalto University School of Business. It comes in subsets by how far the annotators agreed. All of them agreed on 2,264 sentences; at least half agreed on 4,846.

That gap is the first finding. In the broadest subset, 2,582 sentences, or 53%, had at least one annotator who read it differently. A model that "gets one wrong" may have picked a reading a trained person also picked.

Bloomberg's researchers used the same dataset to test BloombergGPT in 2023, and declined to set their scores beside those of custom models that are not large language models, "due to differences in the evaluation setup." The lesson carries over: a score means little outside the setup that produced it.

Three machine readers against the humans

Line up three machine readers against the human labels:

  1. The Loughran-McDonald word lists, built from the full archive of US annual reports plus earnings calls, which count negative, positive and uncertainty words.
  2. A FinBERT model, a language model trained on financial text to label sentiment.
  3. A general chat model, asked to label each sentence with no examples.

Score each on the unanimous subset and on the broad one, build a confusion matrix, then read 20 sentences where machine and humans disagree. Sort them into three piles: the machine was wrong, the label is arguable, or the sentence is genuinely ambiguous. The results table, the matrix and the sorted 20, with a line on each, are the finished project.

ProsusAI's version creates the project's sharpest check. Scored on PhraseBank, it is being tested on the data it learned from, so its accuracy measures memory as much as skill. Saying so in the write-up, or leaving it out of the comparison, is exactly the judgment the project exists to show.

Sentiment analysis also sits in the CFA Institute's Level II practical skills module on Python, data science and AI, which has candidates "explore a common natural language processing task of sentiment analysis."

Test yourself

Partner level

Why is ProsusAI's FinBERT, scored on Financial PhraseBank, a weak test of how well it reads financial sentiment?

Project 4: Judge a fraud model on the frauds it catches

The fourth project leaves language for a classic trap in machine learning: a model that looks excellent on the wrong measure. Here the answer key is certain; what goes wrong is the scoring.

The ULB card-fraud data

The ULB credit card fraud dataset, named for the Université Libre de Bruxelles, holds 284,807 card transactions made by European cardholders over two days in September 2013, collected by Worldline and the university's Machine Learning Group. Only 492 are fraud, 0.172%. For confidentiality, 28 of the features arrive as anonymous principal components; only the time and the amount are left as they were.

The dataset's own page gives the warning in one line: "Given the class imbalance ratio, we recommend measuring the accuracy using the Area Under the Precision-Recall Curve (AUPRC)."

Why 99.8% accuracy can catch nothing

Run the numbers on a model that does nothing. Label every transaction genuine and it is right on 284,315 of 284,807, an accuracy of 99.83%. It catches no fraud at all.

Now take an illustrative model that flags 600 transactions, 400 of them real fraud. Precision, the share of flags that are right, is 400 divided by 600, or 66.7%. Recall, the share of all frauds caught, is 400 divided by 492, or 81.3%. Accuracy rises only to 99.90%.

Two models on the same 284,807 transactionspercent
Do-nothing model, accuracy
99.83%
Do-nothing model, frauds caught
0%
Illustrative model, accuracy
99.90%
Illustrative model, frauds caught
81.3%

Counts from the Worldline and ULB dataset. The second model is illustrative.

Accuracy moved by less than a tenth of a point. The number that matters to a bank went from nothing to four frauds in five.

Validating it the way a bank would

US bank regulators have a formal name for checking a model like this. In April 2026 the Federal Reserve, the OCC and the FDIC replaced SR 11-7, the model risk guidance banks had worked under since 2011, with SR 26-2. It says "Model validation evaluates whether models perform as expected and includes an assessment of a model's reliability and its limitations," and it expects "effective challenge," the critical analysis of a model by "objective experts."

The guidance is written for models like this one: statistical models, and AI that is neither generative nor agentic. Generative and agentic AI models, it says, "are novel and rapidly evolving" and sit outside its scope.

Build the project to that standard, and these four pieces are the finished version:

  • Split by time. Train on the first day and test on the second, which the data allows because its time field counts seconds from the first transaction.
  • Report the right curve. Show the precision-recall curve and its area, not accuracy.
  • Put money on the threshold. Tabulate alerts, frauds caught and money caught at each setting; the amount field lets you weigh the fraud caught against the cost of false alarms.
  • Write the memo. One page on purpose, data limits, performance and known weaknesses, including that the data covers two days in 2013.

Test yourself

Warm-up

On a dataset where 492 of 284,807 card transactions are fraud, what accuracy does a model that flags nothing score?

Project 5: Replicate a published recession model, then try to beat it

The last project turns the checking on yourself. Before claiming a machine learning forecast works, rebuild a model whose answers are already published, and match them. Its trap is time: data that looks clean today is not what anyone knew then.

The New York Fed's one-formula model

In a 2006 paper, Arturo Estrella, then a senior vice president in the New York Fed's research group, and Mary Trubin, who had been an economist in the same group, set out a model that turns the monthly average spread between the 10-year Treasury yield and the three-month Treasury bill into the probability of a US recession 12 months ahead.

Their probit model, estimated on data from January 1959 to December 2005, fits in one Excel formula, with the spread in percentage points in cell A1: =NORMSDIST(-0.6045-0.7374*A1).

Recession probability 12 months ahead, by the Treasury spreadNew York Fed probit, 1959-2005 coefficients
Spread +1.0 point
9.0%
Spread +0.5 point
16.5%
Spread 0
27.3%
Spread -0.5 point
40.7%
Spread -1.0 point
55.3%

Computed from the published coefficients. A negative spread means the yield curve is inverted.

Rebuild it from FRED and expect the first mismatch, because the ready-made series and the paper's inputs are not the same thing.

Beating it without looking ahead

With the benchmark matched, add machine learning: a gradient-boosted model or a regularised regression on FRED-MD, the St. Louis Fed's monthly panel built for big-data macroeconomic research. Test both models on the years after 2005, which the published coefficients never saw. Two traps decide whether the comparison is honest:

  • Revised data. Economic series change after their first release. ALFRED, FRED's archive, lets you "retrieve each economic data release (vintage) that was available on a specific date in history," so the model sees only what forecasters saw.
  • Late recession dates. FRED's recession indicator, USREC, is built from the NBER's business cycle dates, and the NBER names turning points months after they happen.
Turning pointAnnounced by the NBERLag
Peak, December 2007December 1, 2008About 12 months
Trough, June 2009September 20, 2010About 15 months
Peak, February 2020June 8, 2020About 4 months
Trough, April 2020July 19, 2021About 15 months

USREC counts a recession from the first day after a peak, so January 2008 is its first recession month of that cycle. A backtest that treats January 2008 as a known recession month is using information nobody had until December 2008. Label training data with the dates as they were known, or say plainly that you did not.

The finished project puts your replicated coefficients next to the published ones, with every difference explained, and both models' probabilities on one chart with one out-of-sample error measure. It ends in a verdict, even if the verdict is that the 2006 model still wins.

Test yourself

Partner level

Why should a recession backtest not treat each USREC recession month as known at the time it happened?

Which data can you legally use?

Every dataset above is public, but public is not the same as unrestricted, and a project built on data you may not use is a liability rather than a credential. Checking the licence is part of checking the work.

  • Filings. The SEC asks scripts that pull from EDGAR to declare a user agent and keep to at most 10 requests a second. For European companies, filings.xbrl.org, launched by XBRL International in 2021, collects reports prepared in the European Single Electronic Format as they were submitted to national filing mechanisms, with a browsable viewer and a JSON version.
  • FRED. Scripted downloads need a key, and the terms of use warn that some series "may be owned by third parties and subject to copyright restrictions." Using those "for anything other than your own personal use" requires the owner's permission.
  • Research datasets. Financial PhraseBank is licensed CC BY-NC-SA 3.0, the FinanceBench sample CC BY-NC 4.0, and the Loughran-McDonald lists are free for academic research, with a commercial licence on request. All suit a portfolio project; none suits a product you sell.
  • Terminal data. The library guides at NYU and Cornell spell it out. NYU's says data downloaded through the terminal's spreadsheet add-in "should not, under any circumstances, be removed from the Bloomberg terminal," adding: "Only your analysis of the data can be removed." Cornell's says raw data "cannot be removed."

Test yourself

Warm-up

According to university library guides, what may a student take off a Bloomberg Terminal?

Do banks care about AI projects?

Those rules make a project safe to show. Whether a bank wants to see one is a separate question, and banks answer it by describing the skills they want rather than the projects they will ask about. Read what they publish, and each of the five projects produces evidence of it.

WhoWhat they publishWhere and when
JPMorganChase, Data & AI analyst program"the ability to design experiments and deliver measurable outcomes," and to "translate complex technical work for business stakeholders"Careers page, September 2026
UBS, Graduate Talent ProgramAn AI Fluency Pathway "focusing on real-world use cases, responsible application and sound judgment," and advice for applicants: "give us examples and tell us a story"Careers page, September 2026
Morgan Stanley, Code to Give hackathonWinners of its first Montreal event "will be invited to interview for internship opportunities at the firm"Press release, September 2021
CFA InstituteCandidates must complete a practical skills module at each level to receive their exam result; Level II options include Python, data science and AI, at 10 to 20 hours eachProgram page, September 2026

JPMorganChase's wording is the closest thing to a brief: design an experiment, deliver a measurable outcome, then explain it to someone who is not technical. A checking project does all three, and UBS's advice covers how to present it, as an example told as a story. To see how much of the wider ground you already know, the interview readiness check covers the AI tools, roles and bank systems, and scores you as you go.

What a finished project looks like

Each project above ends in its own files. The write-up around them borrows the shape of the validation memo from Project 4, and it is short:

  1. A first sentence that states the finding, with its number.
  2. The data source, its licence and the date you pulled it.
  3. The code or workbook, runnable by someone else from start to finish.
  4. The limits: what the data cannot tell you, and what you would test next.

A reader should get the finding in ten seconds and the method in two minutes. On a CV, the project earns one line: what you tested and what you caught.

Pick the project closest to the seat you want and finish it before starting a second: one finished audit with an honest error rate is worth more than three demos.

The bottom line

The AI projects worth building are not the ones that show a model working. Anyone with an API key can show that. They are the ones that show the model failing on real filings, real transactions and real economic data, and show you catching it: the fiscal year it left unnamed, the frauds it missed while scoring 99.8%, the benchmark it could not beat.

That is the same work JPMorgan's makers-to-checkers shift describes, and what bank regulators call effective challenge. A student who has done it once, on data anyone can download, has something concrete to talk about and a finding to back it.