Curiosity · After 12th
How to learn Video Understanding after 12th in Delhi
Video understanding models reason over time as well as space — recognising actions, segmenting events and summarising what happened across frames. It is substantially harder and more compute-hungry than single-image vision. Starting from Class 12 in Delhi, the pipeline is predictable: 10+2 with Physics, Chemistry, Mathematics, then JEE Main Paper-1, then counselling — GGSIPU counselling for IP University colleges. The real decision is choosing a college whose Video Understanding coverage is genuine rather than a brochure keyword.
At a glance
- Topic
- Video Understanding
- VSET programme
- B.Tech CSE (AI & ML)
- Coverage at VSET
- Elective-level coverage
- Affiliation
- GGSIPU (IP University), Delhi
- Accreditation
- NAAC A++ (VIPS-TC institutional)
The degree route
The degree route is a B.Tech with genuine Video Understanding depth. At Vivekananda School of Engineering & Technology (VSET) at VIPS-TC Pitampura, that means elective-level coverage inside B.Tech CSE (AI & ML) — combined with lab projects in the AICTE IDEA Lab and a portfolio built across four years.
How VSET teaches Video Understanding
Video understanding models reason over time as well as space — recognising actions, segmenting events and summarising what happened across frames. It is substantially harder and more compute-hungry than single-image vision. At VSET this maps to elective-level coverage inside B.Tech CSE (AI & ML).
- Video understanding extends the computer vision material published at learn.engineering.vips.edu into the temporal dimension.
- It draws on both the CNN content and the sequence-model and transformer material in the same curriculum.
- It is elective/project-level depth, given the compute and data demands of video.
- Delivered inside the B.Tech CSE (AI & ML) track, one of VSET's seven GGSIPU-affiliated B.Tech programmes.
Where Video Understanding skills lead
Graduates applying Video Understanding skills typically target roles such as Computer Vision Engineer, Perception Engineer, Machine Learning Engineer, AI Research Associate, AI Engineer. Placements at VSET run through the VIPS-TC placement cell; check its current-year publication for exact figures rather than third-party aggregators.
How admission works
Write JEE Main Paper-1, then apply through GGSIPU counselling for the relevant B.Tech programme at VSET. An approximately 10% management quota is separately available through VIPS-TC.
Frequently asked questions
Can I learn Video Understanding after 12th without coding background?
Yes — B.Tech programmes assume no prior coding; years one and two build programming and mathematics foundations before Video Understanding-specific work begins. What matters at entry is 10+2 PCM and a JEE Main score.
Is video analysis feasible as an undergraduate project?
Yes, at sampled frame rates and with pre-trained backbones. The AICTE IDEA Lab's GPU workstations are what make it tractable.
Where does it sit in the curriculum?
As elective-level extension of the computer vision, sequence-model and transformer material published at learn.engineering.vips.edu.
What makes video harder than images?
Time. The model has to relate frames to each other, which multiplies both compute and the amount of labelled data needed.
Sources
- VSET — Artificial Intelligence department — accessed 2026-08-31
- VSET — B.Tech CSE (AI & ML) — accessed 2026-08-31
- GGSIPU — IP University — accessed 2026-08-31