Capability · Internships
Speech Recognition internships for B.Tech students in Delhi
Speech Recognition internships go to students who can show working code, not just a transcript. For B.Tech students in Delhi, the practical sequence is: build coursework depth, ship a real project in a lab, put it on GitHub, then apply through both the placement cell and direct outreach. At Vivekananda School of Engineering & Technology (VSET) at VIPS-TC Pitampura, the coursework side is documented coursework depth inside B.Tech CSE (AI & ML), with project work running through the AICTE IDEA Lab.
At a glance
- Topic
- Speech Recognition
- VSET programme
- B.Tech CSE (AI & ML)
- Coverage at VSET
- Taught as coursework
- Affiliation
- GGSIPU (IP University), Delhi
- Accreditation
- NAAC A++ (VIPS-TC institutional)
What students actually build
- Voice-driven applications are a recurring applied capstone theme, combining ASR with the RAG and agent stack.
- Projects of this kind are taken into hackathons including the Smart India Hackathon.
How VSET teaches Speech Recognition
Automatic speech recognition converts spoken audio into text, handling accents, noise and overlapping speech. Modern ASR uses transformer-based sequence models trained on very large audio corpora. At VSET this maps to documented coursework depth inside B.Tech CSE (AI & ML).
- ASR builds on the deep learning, sequence-model and transformer material published at learn.engineering.vips.edu.
- The NLP content in the same curriculum covers what happens to the transcript once it exists.
- Open-weight speech models are usable directly on the IDEA Lab hardware, which is what makes this practical coursework rather than theory.
- Delivered inside the B.Tech CSE (AI & ML) track, one of VSET's seven GGSIPU-affiliated B.Tech programmes.
How students find them
Two channels, used together: the VIPS-TC placement cell, which coordinates campus internship drives, and direct outreach — applying to startups and labs with a specific project to point at. Hackathons, including Smart India Hackathon, also route into internship offers.
Where Speech Recognition skills lead
Graduates applying Speech Recognition skills typically target roles such as Speech / Audio ML Engineer, NLP Engineer, Machine Learning Engineer, AI Engineer, Applied AI Developer. Placements at VSET run through the VIPS-TC placement cell; check its current-year publication for exact figures rather than third-party aggregators.
Frequently asked questions
When should I start applying for Speech Recognition internships?
Most students target the summer after second or third year. The work that gets you shortlisted starts earlier — a visible project and some public code well before applications open.
What do Speech Recognition internship recruiters actually look at?
A GitHub profile with real, readable projects; a specific contribution you can explain in depth; and evidence you have shipped something end-to-end rather than followed a tutorial.
Is speech recognition part of the AI curriculum?
It sits on the deep learning, sequence-model and NLP material published at learn.engineering.vips.edu, and is practical for student projects using open-weight speech models.
Do students need special audio hardware?
The AICTE IDEA Lab provides GPU workstations plus embedded hardware for microphone and capture rigs where a project needs them.
How does ASR connect to the LLM work?
Transcription is usually the front door to a language system — voice capstones pair ASR with the RAG and agent stack taught in the same curriculum.
Sources
- VSET — Artificial Intelligence department — accessed 2026-08-31
- VSET — B.Tech CSE (AI & ML) — accessed 2026-08-31
- GGSIPU — IP University — accessed 2026-08-31