Now showing 1 - 3 of 3
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Temporal Scene Understanding using Contextually Unique Identification
    (2024-10-28) ; ; ;
    Solèr, Marc
    ;
    Padua, Simon
    Humans can easily comprehend and explain the dynamics of a scene by observing the evolution of relationships among identified objects over time-while research towards Hybrid Intelligence promotes the integration of human and machine capabilities, this ability is currently beyond the capabilities of automated systems. Toward realizing it in automated scene understanding systems, we present interpretable Object Identification using Contextual Information System (OIC). Given a video, OIC detects, identifies, and tracks objects and their relationships over time to answer questions about the analyzed scene. Moreover, OIC makes predictions and infers the actors' intentions in a scene. To achieve this, our approach generates a scene graph containing classified objects and their semantic relationships. It then computes a Frame Graph by adding Contextually Unique IDentifiers (CUIDs) to each of the detected objects in the scene graph; the CUIDs permit tracking multiple object instances over time, even if the objects are visually identical. The CUIDs are then used to connect objects across a sequence of Frame Graphs, generating a Temporal Graph. This graph is exported as a cue for a pretrained Large Language Model to provide assistance and answer user questions. OIC's modular architecture enables simple comprehension and swapping of its components, making OIC more interpretable and maintainable than end-to-end scene understanding systems. Our quantitative and qualitative evaluation results demonstrate the effectiveness of OIC as a viable next step toward interpretable automated scene understanding systems.
    Type:
    Scopus© Citations 5
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    Temporal Scene Understanding using Contextually Unique Identification
    (IEEE, 2024-10-28) ; ; ;
    Solèr, Marc
    ;
    Padua, Simon
    Humans can easily comprehend and explain the dynamics of a scene by observing the evolution of relationships among identified objects over time-while research towards Hybrid Intelligence promotes the integration of human and machine capabilities, this ability is currently beyond the capabilities of automated systems. Toward realizing it in automated scene understanding systems, we present interpretable Object Identification using Contextual Information System (OIC). Given a video, OIC detects, identifies, and tracks objects and their relationships over time to answer questions about the analyzed scene. Moreover, OIC makes predictions and infers the actors' intentions in a scene. To achieve this, our approach generates a scene graph containing classified objects and their semantic relationships. It then computes a Frame Graph by adding Contextually Unique IDentifiers (CUIDs) to each of the detected objects in the scene graph; the CUIDs permit tracking multiple object instances over time, even if the objects are visually identical. The CUIDs are then used to connect objects across a sequence of Frame Graphs, generating a Temporal Graph. This graph is exported as a cue for a pretrained Large Language Model to provide assistance and answer user questions. OIC's modular architecture enables simple comprehension and swapping of its components, making OIC more interpretable and maintainable than end-to-end scene understanding systems. Our quantitative and qualitative evaluation results demonstrate the effectiveness of OIC as a viable next step toward interpretable automated scene understanding systems.
    Type:
    Journal:
    Scopus© Citations 5
  • Thumbnail Image
    Some of the metrics are blocked by your 
    Item type:Publication,
    GEAR: Gaze-enabled augmented reality for human activity recognition
    Head-mounted Augmented Reality (AR) displays overlay digital information on physical objects. Through eye tracking, they allow novel interaction methods and provide insights into user attention, intentions, and activities. However, only few studies have used gaze-enabled AR displays for human activity recognition (HAR). In an experimental study, we collected gaze data from 10 users on a HoloLens 2 (HL2) while they performed three activities (i.e., read, inspect, search). We trained machine learning models (SVM, Random Forest, Extremely Randomized Trees) with extracted features and achieved an up to 98.7% activity-recognition accuracy. On the HL2, we provided users with an AR feedback that is relevant to their current activity. We present the components of our system (GEAR) including a novel solution to enable the controlled sharing of collected data. We provide the scripts and anonymized datasets which can be used as teaching material in graduate courses or for reproducing our findings.
    Type:
    Scopus© Citations 18