A web platform that uses artificial intelligence to read scanned documents, spreadsheets, and historical archives, helping officials generate geological and mining reports automatically.
Ministry of Coal · Coal India Limited · Software
Signing in saves it for your whole team — everyone on your invite link sees the same two entries. Anything you shortlisted while signed out comes with you.
Decomposed from what the description asks for. Nothing added.
Document Ingestion
A module to upload and process scanned PDFs, digital documents, spreadsheets, and images from historical archives.
Automated Report Generator
A system component that compiles data and builds structured production and geological reports automatically.
Topic Identification Module
A feature that performs automated word cloud generation and topic identification on the ingested text.
Query and Response System
An artificial intelligence-based assistant that lets users ask questions and retrieve insights from the datasets quickly.
Both columns are read off the brief's own wording. Nothing here is inferred from the ministry's name.
The jury will check whether your platform can handle multiple file formats like scanned PDFs and spreadsheets, accurately extract data, and generate reports or answers quickly to reduce manual effort.
A jury can still ask about these. Decide them deliberately rather than by accident.
Generated from the brief's own wording and the competition's published rules — never from a guess about what this ministry prefers.
Exact templates required for the final reports?
The brief never answers this, so a panel will. Whatever you decide, say it the same way twice.
Specific target percentages for accuracy and reduction in report preparation time?
The brief never answers this, so a panel will. Whatever you decide, say it the same way twice.
Sample historical datasets from CMPDI or CIL subsidiaries?
The brief never answers this, so a panel will. Whatever you decide, say it the same way twice.
Has any part of this been shown at a previous event, hackathon or college project?
The guidelines are explicit: your solution must not have appeared in any previous event or programme, of any sort. A recycled project is what a team under time pressure reaches for.
Each one is quoted from a gap in the brief, not a guess about your team.
Needs data you may not get
The problem statement relies on historical archives, scanned PDFs, and mining figures from CMPDI and CIL subsidiaries which are not provided to students.
No measurable target
The description asks for percentage reductions and maximum accuracy calculations without providing exact baseline numbers or target thresholds.
What the organisers attached, and what the brief assumes you can get.
Same organisation, same year. Reading two of theirs tells you more about what they care about than reading one.
Pick what you are about to do and copy the prompt. It carries the organisers' own wording, the constraints they never spell out, and an instruction not to invent requirements they never set.
Who has this problem, what already exists, and what you would have to find out.
The brief asks for document ingestion. How would you build that?
Decoded from SIH26023 itself — A module to upload and process scanned PDFs, digital documents, spreadsheets, and images from historical archives. The brief asks for it by name.
The brief asks for automated report generator. How would you build that?
Decoded from SIH26023 itself — A system component that compiles data and builds structured production and geological reports automatically. The brief asks for it by name.
The brief asks for topic identification module. How would you build that?
Decoded from SIH26023 itself — A feature that performs automated word cloud generation and topic identification on the ingested text. The brief asks for it by name.
The brief asks for query and response system. How would you build that?
Decoded from SIH26023 itself — An artificial intelligence-based assistant that lets users ask questions and retrieve insights from the datasets quickly. The brief asks for it by name.
Where does your data come from — a published source, one you collect, or one you generate?
No dataset is attached to this problem statement, so sourcing it is part of the work and nobody told you that.
Why not use what already exists? Name the closest thing to this that is already running.
A team that has not named the alternative themselves is answering this for the first time in the room.
Which single thing will you demonstrate end to end, start to finish, with nothing skipped?
Ours, not a rule: a narrow thing that fully works survives questioning better than a broad thing that half works. If nobody on the team can name it, that is the finding.
Show me this working: the jury will check whether your platform can handle multiple file formats like scanned PDFs and spreadsheets, accurately extract data, and generate reports or answers quickly to reduce manual effort.
This is the evaluator read for your problem statement, decoded from the brief's own wording.