COMP 648: Computer Vision Seminar | Fall 2026

Instructor: Vicente Ordóñez-Román (vicenteor at rice.edu)
Class Time: Tuesdays from 4pm to 5:15pm Central Time (Howard Keck Hall, Room 107).

Course Description: This seminar will explore and analyze the current literature in computer vision, especially focusing on computational methods for visual recognition. Our topics include image classification and understanding, object detection, image segmentation, and other high-level perceptual tasks. We will explore this semester recent vision foundation models, and multimodal foundation models that involve images, video and language. This is a 1-credit graduate seminar with student-led weekly presentations.

Prerrequisite: COMP 646 (Deep Learning for Vision and Language) or research experience in deep learning, computer vision or related fields. Ask the instructor if you are unsure.

Schedule

Date Topic  
Aug 25th Welcome & Overview
  • Introductions and an overview of the topics to be covered this semester, along with the schedule of student-led presentations.
Sep 1st
Steerable Visual Representations. April 2026. [technical report]
  • Injects text into the layers of a ViT via lightweight cross-attention, so its global and local features can be steered with language rather than always latching onto the most salient object.
  • Presentation led by Jefferson Hernandez
Sep 8th
TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration. June 2026. [technical report]
  • Introduces a template-guided iterative framework that proactively discovers multiple hidden problems in personal workspaces and software repositories, grounds them in contextual evidence, and proposes concrete actions.
  • Presentation led by Jaywon Koo
Sep 15th
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence. August 2026. [technical report]
  • Stress-tests LLM judges for stability rather than accuracy. Across 9 frontier models and 14 tasks, verdicts flip 25–71% of the time under pushback, and 62–91% against an adversarial persuader.
  • Presentation led by Ami Aihara
Sep 22nd
Revisiting Autoregressive Models for Generative Image Classification. March 2026. [technical report]
  • Class-conditional generative models can serve as classifiers, but autoregressive ones have been held back by a fixed token order that biases what the model attends to. Marginalizing predictions over many orders with any-order AR models beats diffusion-based classifiers while being up to 25x more efficient, and comes close to self-supervised discriminative models.
  • Presentation led by Catherine He
Sep 29th
Controlling Embedding Spaces with Text-Conditioned Transformations. July 2026. [technical report] [project page]
  • Learns a text-conditioned affine transformation of frozen CLIP embeddings to surface attributes a single vector suppresses, such as color or camera angle, steering the embedding space itself rather than the encoder, with no re-encoding needed.
  • Presentation led by TBD
Oct 6th
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance. April 2026. [technical report]
  • VLMs recognize semantics well but lack spatial invariance, misjudging object identity under simple rotations and scaling. The failure is sharpest when semantic content is sparse, and holds across architectures and model sizes.
  • Presentation led by TBD
Oct 13th MIDTERM RECESS (NO SCHEDULED CLASSES)
Oct 20th
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus. August 2026. [technical report]
  • Verifies each step of a multimodal reasoning chain by coupling several frozen, training-free verifiers and treating their disagreement as signal rather than noise, solved in closed form as a coordination game.
  • Presentation led by TBD
Oct 27th
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning. April 2026. [technical report]
  • Decomposes text-to-image synthesis into an interleaved trajectory of textual planning, visual drafting, textual reflection, and visual refinement, with dense supervision keeping the intermediate states spatially and semantically consistent.
  • Presentation led by TBD
Nov 3rd
Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs. April 2026. [technical report] [project page]
  • Curates a dictionary of 15,000 multimodal concepts from over 400,000 caption-image pairs, then trains a sparse autoencoder aligned to that dictionary to steer a frozen MLLM's activations at inference time. Gives granular, concept-specific safety control without retraining, improving robustness on benchmarks like MM-SafetyBench and JailBreakV while preserving general-purpose capability.
  • Presentation led by TBD
Nov 10th Tentative topic: Multimodal Agents. Paper TBD.
Nov 17th
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning. March 2026. [technical report] [code] [project page]
  • Sign language production is caught between text-to-pose models that regress to the mean and dictionary retrieval that stitches together disjointed transitions. Learning from sparse keyframes captures the kinematic distribution of signing instead, with a conditional flow matching framework spanning four sign languages and photorealistic rendering through 3D Gaussian splatting.
  • Presentation led by TBD
Nov 24th Tentative topic: 3D and Embodied Visual Reasoning. Paper TBD.
Dec 1st Tentative topic: Emerging Directions in Multimodal Foundation Models. Paper TBD.

Disclaimer: The topics on this list are tentative and subject to adjustments throughout the semester as interests in the group evolve.

Logistics: This is a seminar with a pass/fail grade. Registered students are required to participate and present a recent work in a topic of interest of the seminar at least once throughout the semester. A Satisfactory grade requires participating presenting a paper at least once during the semester and actively participating in discussions throughout the semester. Students registered for 3 credits must additionally complete a project, including a project report. The scope and extent of the project report should be agreed upon between the student and the instructor.

Honor Code and Academic Integrity: "In this course, all students will be held to the standards of the Rice Honor Code, a code that you pledged to honor when you matriculated at this institution. If you are unfamiliar with the details of this code and how it is administered, you should consult the Honor System Handbook at http://honor.rice.edu/honor-system-handbook/. This handbook outlines the University's expectations for the integrity of your academic work, the procedures for resolving alleged violations of those expectations, and the rights and responsibilities of students and faculty members throughout the process."

Title IX Support: Rice University cares about your wellbeing and safety. Rice encourages any student who has experienced an incident of harassment, pregnancy discrimination or gender discrimination or relationship, sexual, or other forms interpersonal violence to seek support through The SAFE Office. Students should be aware when seeking support on campus that most employees, including myself, as the instructor/TA, are required by Title IX to disclose all incidents of non-consensual interpersonal behaviors to Title IX professionals on campus who can act to support that student and meet their needs. For more information, please visit safe.rice.edu or email titleixsupport@rice.edu.

Disability Resource Center: "If you have a documented disability or other condition that may affect academic performance you should: 1) make sure this documentation is on file with the Disability Resource Center (Allen Center, Room 111 / adarice@rice.edu / x5841) to determine the accommodations you need; and 2) talk with me to discuss your accommodation needs."