As cognitive abilities decline, even familiar tasks can become difficult to do without help. Researchers at the University of Bristol are using supercomputing to process thousands of hours of video and develop AI models that understand how a person performs tasks, anticipate their actions, and flag unusual behaviour. The research could eventually help people experiencing cognitive decline live independently for longer and reduce their need for constant in-person care. Importantly, it may ease the pressure on caregivers and healthcare systems as populations age.
Tracking routines to support independent living
When cognitive abilities start to decline, ordinary tasks such as making coffee or taking medication can become difficult to do. Small lapses in routine can slowly erode someone’s independence and increase their need for daily support. While existing assistive tools help, they often fail the moment a person pauses or forgets what they’re supposed to do.
Researchers at the University of Bristol are taking a different approach to assisted living. Using the Isambard-AI supercomputer at the Bristol Centre for Supercomputing, the team has processed more than 3,600 hours of video recorded globally from body-worn cameras and developed models that understand how a person performs tasks and anticipates their actions.
This kind of predictive assistance technology could support wearable tools such as smart glasses that offer prompts when someone forgets a step or needs support completing a task, helping them live independently for longer.
“An assistive system built on this research will be able to remember what you’ve done,” says Dr Michael Wray, a member of the team and Assistant Professor in Computer Vision at the University of Bristol’s School of Computer Science. “It will be able to work out if you’re doing things differently to normal.”
Understanding how people do their tasks
Many assistive tools can identify basic activities such as whether someone is walking or sitting. However, they struggle to follow the sequences of smaller actions that make up daily tasks, especially when a person improvises or pauses midway.
“If we know what a person does in their daily routine of tasks, we can anticipate what they might do next,” explains Wray. “And then if they deviate from that, maybe it’s clear that they’re about to fall over or about to touch something hot, then we can intervene.”
The research focuses on egocentric vision, with participants wearing head-mounted cameras while going about their usual routines, particularly in kitchens. Kitchens pose a unique set of challenges due to the high volume of tasks and potentially dangerous objects.
This method captures small differences in how people do tasks, which standard datasets often miss. “The way I cut onions is going to be different to someone else and how they cut onions,” says Wray, adding that even placing a spoon on a counter can have meaning that only makes sense based on someone’s habits.
“To the individual doing it, it’s like, ‘Oh, of course I’m putting my spoon down here because I might need it later’. But to someone watching, they might think, ‘But why would you put that spoon down there and not on the spoon rest?’”
Developing systems that understand context
While earlier AI models often treated video as disconnected images, the team’s research focuses on how actions connect from one moment to the next.
“If in one frame you’re letting go of a cup, the next frame is going to show the cup start to fall, and then a few frames after that, the cup is going to smash on the ground,” notes Wray. “But early video models weren’t operating in a way that made sense of this fact, or understood that time goes in one direction.”
The team’s models analyse short sequences of video and use those recordings to build digital twins that recreate a person’s surroundings in 3D. As Wray explains, “Having a digital twin lets us know where objects are being placed, and we can get feedback on where people are looking when they’re placing things.”
Eye tracking also gives researchers clues about someone’s intent before they act. Humans naturally look where they’re about to place something a second or two before they do it. Capturing eye movement can help AI tools anticipate more accurately what someone is about to do.
This level of understanding could help people retrace their misplaced items or remember what they were supposed to do.
Training models on much larger video datasets
With thousands of hours of video to process, the team’s previous systems were pushed to their limits. “It was either incredibly slow or we just couldn’t use it at all,” says Wray. “We had to use smaller partitions of the datasets. We were very limited in the scope of projects that we could look into.”
The team has since been able to process much larger datasets and train more complex models with the rollout of Isambard-AI, a supercomputer hosted at the Bristol Centre for Supercomputing. “The scale that Isambard-AI gives us is completely different to what we had access to before, and it unlocks a whole new set of tasks for us,” shares Wray. “It allows us to train a lot more models and to test a lot more things.”
Isambard-AI is an HPE Cray Supercomputing EX system with more than 5,000 NVIDIA® GH200 Grace Hopper Superchips linked by HPE Slingshot 11. It was built entirely within a dedicated and energy-efficient AI Mod POD from HPE.
Greater compute capacity through Isambard-AI enables researchers to work with full datasets instead of just small subsets, making their findings more applicable to real-life environments where variability can’t be avoided. It also allows them to study human behaviour over longer periods of time.
Easing the load of families and caregivers
Although this type of predictive assistance technology is still in the early stages of being researched and hasn’t yet been tested in real-world environments, Wray believes it could help people maintain their independence longer while reducing pressure on families, caregivers, and healthcare systems.
“Our aim is to develop the technology so it can provide a high level of care without as much human intervention,” he says. “It would free up resources from the National Health Service or carers to really help out with an aging population, like we have in the UK and other parts of the world.”
With birth rates declining across the globe, Wray says it will be difficult to maintain a high standard of care by just relying on more people. Predictive assistance technology could help narrow this gap and reduce the reliance on caregivers.
“Maybe we could cut their visits down to every couple of days,” he adds. “Or instead of twice a day, it could be just once a day, because you have this device which can provide a lot of that support autonomously.”
Extending into robotics and industrial environments
Outside of assisted living, the research could be used in robotics since robots run into many of the same problems that humans do when moving through unfamiliar spaces and judging how to interact with their surroundings.
Wray cites earlier research where systems trained on the team’s first-person video dataset later performed better in robotics experiments, even in completely different environments. The systems got better at tasks such as stacking objects and locating items on cluttered desks.
The technology may also have potential in industrial environments. Tasks in these environments are often more repetitive and predictable than they are in homes, making it easier for systems to recognize when something is wrong or when someone is struggling with a task.
“If you know specifically what’s meant to be happening, then it can be relatively easy to find these kinds of anomalies,” says Wray. “Anything that goes against the grain of what you might expect stands out more clearly. This technology could be a huge help there.”
For now, the team is focused on solving the challenges of tracking how people do routine tasks over longer periods of time. “If you want something that you put on and then three months later it can still be helping you and keeping track of what’s been going on during those three months, that’s quite far beyond current technology,” says Wray.