Information Asymmetry in Indian Derivatives Markets
Three months on a study of information asymmetry in Indian derivatives markets. I audited all 46 references against the primary sources after finding I could not defend the paper. It is paused until I understand its statistics well enough to defend those too.
Can I actually defend this, or does it only look finished?
This one is unfinished. I have put it first because the most useful thing it produced was not a result, it was finding out that I could not defend my own paper.
What I was trying to do
India’s National Stock Exchange publishes something most exchanges do not. Every day it releases derivatives trading volume broken down by participant category: foreign institutions, domestic institutions, proprietary desks and retail clients, each reported separately. Almost nowhere else can you see who is on which side of the market.
I wanted to know whether that data showed institutions knowing things retail traders did not. The test was direct. Build a directional positioning measure for each category from their options and futures activity, then check whether it predicted the next day’s Nifty 50 return. I ran it across 1,032 trading days from January 2022 to March 2026, extending a method Chang et al. (2009) had applied to Taiwan’s options market.
It produced results. Institutional options positioning predicted the next day’s return, retail positioning predicted it in the opposite direction, the effect sat in options rather than futures, and it faded over the following week. It survived controls for momentum and volatility.
Where it went wrong
I had a document that looked like a paper. Abstract, literature review, six hypotheses, regression tables, forty-six references. Then someone asked me a question about it and I could not answer. I knew where each sentence had come from. I could not say why any of it was true.
I had assembled the literature review with AI help and never checked the output against the sources.
Checking the sources
I went through all 46 references. For each one I found every place it was used in the paper, wrote down the exact sentence, and read the source to see whether it supported that sentence. Not whether the reference existed. Whether the claim in front of it was true.
The errors were not the kind I had expected. One paper was cited to support a statement it argues against: Hong and Stein (1999) appeared behind a claim that trend-following rules have little predictive power, when their model is built on the premise that such strategies do work in the short run. An argument was attributed to Hasbrouck (1995) that belongs to Chan (1992) and Stoll and Whaley (1990). A finding from Chang et al. (2009) was stated backwards. Easley, O’Hara and Srinivas (1998) was credited with a claim about futures markets, and that paper does not discuss futures markets.
Those are much harder to catch than an invented reference. The paper is real, the authors are right, the year is right, and the sentence in front of the citation is still wrong. A reader checking whether the sources exist finds nothing wrong at all.
It took weeks and it was extremely boring. About twenty fixes came out of it across fourteen drafts.
Where it stands
The references are repaired. All forty-six have been checked against the primary sources, and the claims in the text now match what those sources actually say.
The statistics are a different problem, and they are why it is still paused. The methods in the paper are more rigorous than anything I could have reached on my own: Newey-West standard errors, stationarity testing, regime conditioning, corrections for multiple comparisons. An AI put them there and they made the paper genuinely stronger. I do not understand them well enough to defend them if somebody pressed me on why each one was the right choice.
I could publish it anyway. Nobody would stop me. But then it would be a paper an AI thought and I typed, and I want this one to be mine. If the machine does the thinking and applies the statistics, I am not sure what I am for.
So it waits until I am competent enough in statistics to stand behind every line of it.
What changed afterwards
I now check whether a cited source says what the sentence citing it claims, instead of treating a bibliography as evidence by itself. Before showing anyone anything, I go looking for the part I would have trouble explaining, and I start there. I have also stopped reading fluent, confident writing as a sign that something is correct, whether a model produced it or I did.
- Research
- Market microstructure
- Method