Hardik Gulati
Third-year B.Tech student in Artificial Intelligence & Machine Learning with a strong foundation in Python, data analysis, and statistics. I build practical AI and data-driven systems — recent work spans a retrieval-augmented research assistant, an accessibility automation tool, and a study of text-extraction errors in government PDFs. Interested in roles across machine learning, backend systems, data, and problem-solving.
audited
variants found
shipped
deployments
Selected Work
2025 — present-
A11y GitHub Issue Logger
A tool that scans a public URL with axe-core, checks the target repository for duplicate issues, drafts a structured WCAG accessibility issue, and logs it to GitHub after review.
- Bring-your-own-token model, so no user credentials are ever stored server-side.
- GitHub Actions CI runs lint, type checks, and a pytest/Playwright suite on every pull request.
- Backend deployed to Railway, frontend to Vercel; API rate-limited with slowapi.
- Built across feature branches with reviewed pull requests.
FastAPIReactPlaywrightaxe-corepytestGitHub ActionsJul 2025 —
Present -
Scholar — Retrieval-Augmented Research Assistant
A document Q&A system that ingests PDF, DOCX, and PPTX files, chunks and embeds them with sentence-transformers, and answers questions using FAISS vector search combined with BM25 keyword search.
- Citation layer points every answer back to the source passage it came from.
- An evaluation script measures retrieval recall on a test question set — retrieval configurations are compared, not assumed.
- Data model in SQLModel/SQLite; background jobs handled with arq.
FastAPIReactFAISSsentence-transformersBM25SQLModelarqPresent -
DevAudit / SILT — Text-Layer Corruption in Government PDFs
Measuring how often born-digital Indian government PDFs carry a Devanagari text layer that extracts without error but is silently wrong — a failure caused by legacy non-Unicode fonts.
1,602 documents 8 issuing bodies 3 states 141 legacy font variants- Five-bucket classification applied across the full corpus.
- Decision threshold and random seed pre-registered before running the audit.
PythonPDF parsingStatisticsPresent
Earlier Projects
Evidence-grounded chest X-ray reporting pipeline that verifies its own clinical claims — ViT + DistilGPT-2 generation with FAISS retrieval and sentence-level confidence filtering.
Banking transaction analyser scoring payments against Logistic Regression and Random Forest pipelines, served by a Spring Boot API over MySQL with a JavaFX desktop client.
Donation rerouting platform for orphanages, redistributing surplus resources to the shelters that need them most. Role-based auth, pledge tracking, and SQL-driven analytics.
Financial document processing for Indian MSMEs — extracts vendor, amount, GST, date, and invoice number from photos or PDFs, categorises the expense, and feeds a spending dashboard.
Skills
Education & Credentials
B.Tech, Artificial Intelligence & Machine Learning
Class XII — CBSE
Class X — CBSE
Technical Volunteer
Deep Learning
Supervised Machine Learning
Artificial Intelligence
Hashgraph Developer
Machine Learning Specialization
Let's build
something real.
Looking for roles and internships across machine learning, backend systems, and data. If you have a problem worth solving, I'd like to hear about it.