My cousin studies Economics at college, and she asked me whether she should consider doing a Ph.D. given what AI models can do now.
Although I didn’t study the subject in college and only know the basic concepts, I shared my thoughts and told her it was strictly my opinion. I also told her to seek more opinions from people who are currently getting a Ph.D.
As soon as our conversation ended, I wondered how well today’s AI models can handle economics.
I decided to find out. I picked the Reserve Bank of India’s (RBI’s) long Annual Report document and uploaded it to the three leading AI models.
I was not trying to find out whether an AI model could make a Ph.D. in economics obsolete or less relevant. Instead, I wanted to see how well these models could understand a long, information-heavy economics document.
I tested three leading AI models on a 200-page document and found which one actually comprehends texts that would take humans hours.
I tested GPT-5.6 Luna, Claude’s Sonnet 5, and Gemini 3.6 Flash
I could easily pick the winner
One of the key components of the RBI’s Annual Report is an assessment of India’s economic and financial performance. I’m most worried about inflation, so I usually read that part carefully every year.
My questions focused on inflation. I asked all the models the same question to compare their answers fairly.
I’d start with what I felt was the worst response. It came from Gemini 3.6 Flash. Of the three, its response felt the most mechanical, like I got a response from a robot.
While I liked Gemini’s short response, it was formatted in a way that made even the short length feel intimidating to read.
On the bright side, it analyzed the document well and didn’t hallucinate while answering the questions. It also gave a nice summary of the importance towards the end, but who would reach that point when the rest of the answer feels intimidating to read?
Gemini 3.6 Flash was better than GPT-5.6 Luna and Sonnet 5. It was hard for me to pick one of these two as the winner, but I ultimately chose Sonnet 5 because it felt more human than GPT-5.6 Luna.
In its short intro, Sonnet 5 highlights the section it analyzed to get you the answer so that you can fact-check it quickly. This is consistent with all its answers, and I absolutely love it.
ChatGPT gets straight to the point, which isn’t bad either, but it felt slightly more mechanical than Claude’s answer. However, I’d have to give it to GPT-5.6 Luna for a more structured answer using a table.
I also set a false premise and asked a question to see how these models react to a wrong fact. All of them spotted that I reversed the premise and highlighted the facts the RBI posted in its Annual Report. Again, I loved Sonnet 5’s framing of the answer the most.
All of them were also good at interpreting data given in the document, but I kept Sonnet 5 ahead because of its easy-to-read answer.
I asked many more questions and observed the same trend throughout my session, which was enough for me to pick Sonnet 5 as the winner.
Gemini 3.6 Flash and GPT-5.6 Luna aren’t bad
They aren’t just suitable for my taste
I heard complaints on Reddit that Gemini only analyzes the first and last page of a long PDF and hallucinates a lot. This isn’t my experience with Google’s AI model. Instead, I found it answered all my questions correctly.
Although it sounds more mechanical, the crisp summary toward the end helps if you don’t want to read its boring answer. Based on my experience, I feel the Gemini AI model is even better at summarizing.
If Gemini 3.6 Flash and Sonnet 5 are the two extremes, then GPT-5.6 Luna sits right in the middle. My testing suggests that it has qualities from both and also offers structured data formatting in its answers more often than others.
When I want to read something important, I would prefer an answer that gives me everything I need to know in an easy-to-understand way. I wouldn’t mind the absence of a table in the answer if the text is easy to read.
This test reinforced something important about AI models
GPT-5.6 Luna has a context window of 1,050,000 tokens, while Sonnet 5 has 1,000,000. At 1,048,576, Gemini 3.6 Flash has slightly more than Sonnet 5.
If I go purely by numbers, Sonnet 5 is the least capable. However, the results paint a completely different picture. A larger context window only means it can hold more data and doesn’t guarantee it’ll process information more accurately.
Also, longer context can lead to poorer reasoning quality. While more context certainly helps, too much of it can make it harder for the Large Language Models (LLMs) to answer accurately.
The goal should always be striking that sweet spot between giving AI enough context and giving it too much.


