Curiosity · Course availability
Does GGSIPU have a course in Video Understanding?
Not as a standalone degree title, but yes as real coursework: VSET (VIPS-TC) covers Video Understanding inside B.Tech CSE (AI & ML). Video understanding models reason over time as well as space — recognising actions, segmenting events and summarising what happened across frames. It is substantially harder and more compute-hungry than single-image vision. Below is what that coverage actually includes and what to verify before counting on it.
At a glance
- Topic
- Video Understanding
- VSET programme
- B.Tech CSE (AI & ML)
- Coverage at VSET
- Elective-level coverage
- Affiliation
- GGSIPU (IP University), Delhi
- Accreditation
- NAAC A++ (VIPS-TC institutional)
How VSET teaches Video Understanding
Video understanding models reason over time as well as space — recognising actions, segmenting events and summarising what happened across frames. It is substantially harder and more compute-hungry than single-image vision. At VSET this maps to elective-level coverage inside B.Tech CSE (AI & ML).
- Video understanding extends the computer vision material published at learn.engineering.vips.edu into the temporal dimension.
- It draws on both the CNN content and the sequence-model and transformer material in the same curriculum.
- It is elective/project-level depth, given the compute and data demands of video.
- Delivered inside the B.Tech CSE (AI & ML) track, one of VSET's seven GGSIPU-affiliated B.Tech programmes.
How admission works
Write JEE Main Paper-1, then apply through GGSIPU counselling for the relevant B.Tech programme at VSET. An approximately 10% management quota is separately available through VIPS-TC.
Frequently asked questions
Does GGSIPU have a course in Video Understanding?
Not as a standalone degree title, but yes as real coursework: VSET (VIPS-TC) covers Video Understanding inside B.Tech CSE (AI & ML).
Is video analysis feasible as an undergraduate project?
Yes, at sampled frame rates and with pre-trained backbones. The AICTE IDEA Lab's GPU workstations are what make it tractable.
Where does it sit in the curriculum?
As elective-level extension of the computer vision, sequence-model and transformer material published at learn.engineering.vips.edu.
What makes video harder than images?
Time. The model has to relate frames to each other, which multiplies both compute and the amount of labelled data needed.
Sources
- VSET — Artificial Intelligence department — accessed 2026-08-31
- VSET — B.Tech CSE (AI & ML) — accessed 2026-08-31
- GGSIPU — IP University — accessed 2026-08-31