
Laparoscopy is a surgical technique in which a fiber-optic camera is inserted into a patient’s abdominal cavity to provide a video feed that guides the surgeon through a minimally invasive procedure. Laparoscopic surgeries can take hours, and the video generated by the camera — the laparoscope — is often recorded. Those recordings contain a wealth of information that could be useful for training both medical providers and computer systems that would aid with surgery, but because reviewing them is so time consuming, they mostly sit idle.
Researchers at MIT and Massachusetts General Hospital hope to change that, with a new system that can efficiently search through hundreds of hours of video for events and visual features that correspond to a few training examples.
In work they presented at the International Conference on Robotics and Automation this month, the researchers trained their system to recognize different stages of an operation, such as biopsy, tissue removal, stapling, and wound cleansing.
But the system could be applied to any analytical question that doctors deem worthwhile. It could, for instance, be trained to predict when particular medical instruments — such as additional staple cartridges — should be prepared for the surgeon’s use, or it could sound an alert if a surgeon encounters rare, aberrant anatomy.
“Surgeons are thrilled by all the features that our work enables,” says Daniela Rus, an Andrew and Erna Viterbi Professor of Electrical Engineering and Computer Science and senior author on the paper. “They are thrilled to have the surgical tapes automatically segmented and indexed, because now those tapes can be used for training. If we want to learn about phase two of a surgery, we know exactly where to go to look for that segment. We don’t have to watch every minute before that. The other thing that is extraordinarily exciting to the surgeons is that in the future, we should be able to monitor the progression of the operation in real-time.”
Joining Rus on the paper are first author Mikhail Volkov, who was a postdoc in Rus’ group when the work was done and is now a quantitative analyst at SMBC Nikko Securities in Tokyo; Guy Rosman, another postdoc in Rus’ group; and Daniel Hashimoto and Ozanan Meireles of Massachusetts General Hospital (MGH).
Representative frames
The new paper builds on previous work from Rus’ group on “coresets,” or subsets of much larger data sets that preserve their salient statistical characteristics. In the past, Rus’ group has used coresets to perform tasks such as deducing the topics of Wikipedia articles or recording the routes traversed by GPS-connected cars.
In this case, the coreset consists of a couple hundred or so short segments of video — just a few frames each. Each segment is selected because it offers a good approximation of the dozens or even hundreds of frames surrounding it. The coreset thus winnows a video file down to only about one-tenth its initial size, while still preserving most of its vital information.
For this research, MGH surgeons identified seven distinct stages in a procedure for removing part of the stomach, and the researchers tagged the beginnings of each stage in eight laparoscopic videos. Those videos were used to train a machine-learning system, which was in turn applied to the coresets of four laparoscopic videos it hadn’t previously seen. For each short video snippet in the coresets, the system was able to assign it to the correct stage of surgery with 93 percent accuracy.
“We wanted to see how this system works for relatively small training sets,” Rosman explains. “If you’re in a specific hospital, and you’re interested in a specific surgery type, or even more important, a specific variant of a surgery — all the surgeries where this or that happened — you may not have a lot of examples.”
Selection criteria
The general procedure that the researchers used to extract the coresets is one they’ve previously described, but coreset selection always hinges on specific properties of the data it’s being applied to. The data included in the coreset — here, frames of video — must approximate the data being left out, and the degree of approximation is measured differently for different types of data.
Machine learning can be thought of as a problem of approximation, however. In this case, the system had to learn to identify similarities between frames of video in separate laparoscopic feeds that denoted the same phases of a surgical procedure. The metric of similarity that it arrived at also served to assess the similarity of video frames that were included in the coreset, to those that were omitted.
“Interventional medicine — surgery in particular — really comes down to human performance in many ways,” says Gregory Hager, a professor of computer science at Johns Hopkins University who investigates medical applications of computer and robotic technologies. “As in many other areas of human endeavor, like sports, the quality of the human performance determines the quality of the outcome that you achieve, but we don’t know a lot about, if you will, the analytics of what creates a good surgeon. Work like what Daniela is doing and our work really goes to the question of: Can we start to quantify what the process in surgery is, and then within that process, can we develop measures where we can relate human performance to the quality of care that a patient receives?”
“Right now, efficiency” — of the kind provided by coresets — “is probably not that important, because we’re dealing with small numbers of these things,” Hager adds. “But you could imagine that, if you started to record every surgery that’s performed — we’re talking tens of millions of procedures in the U.S. alone — now it starts to be interesting to think about efficiency.”











The second panel featured talks on sector based solutions starting with the International Federation of the Red Cross (IFRC). The Federation (Aarathi) spoke about their joint project with WeRobotics; looking at cross-sectoral needs for various robotics solutions in the South Pacific. IFRC is exploring at the possibility of launching a South Pacific Flying Labs with a strong focus on women and girls. Pix4D (Lorenzo) addressed the role of aerial robotics in agriculture, giving concrete examples of successful applications while providing guidance to our Flying Labs Coordinators.
Panel number three addressed the transformation of transportation. UNICEF (Judith) highlighted the field tests they have been carrying out in Malawi; using cargo robotics to transport HIV samples in order to accelerate HIV testing and thus treatment. UNICEF has also launched an air corridor in Malawi to enable further field-testing of flying robots. MSF (Oriol) shared their approach to cargo delivery using aerial robotics. They shared examples from Papua New Guinea












I enjoyed talking at the Visual Localization session (another packed out session with standing room only) on our paper, “
The social functions were good and the conference hotel was incredible, especially the infinity pool at the top of the hotel, 57 floors up, looking over the Marina. The night safari was fun too.
Overall, a great week and always great to reconnect with hundreds of colleagues and collaborators from around the world. See you next year in Brisbane!
Mosquitos kill more humans every year than any other animal on the planet and conventional methods to reduce mosquito-borne illnesses haven’t worked as well as many hoped. So we’ve been hard at work since receiving this USAID grant six months ago to reduce Zika incidence and related threats to public health.
Our approach seeks to complement and extend (not replace) these existing delivery methods. The challenge with manned aircraft is that they are expensive to operate and maintain. They may also not be able to target areas with great accuracy given the altitudes they have to fly at.
Cars are less expensive, but they rely on ground infrastructure. This can be a challenge in some corners of the world when roads become unusable due to rainy seasons or natural disasters. What’s more, not everyone lives on or even close to a road.
Our IAEA colleagues thus envision establishing small mosquito breeding labs in strategic regions in order to release sterilized male mosquitos and reduce the overall mosquito population in select hotspots. The idea would be to use both ground and aerial release methods with cars and flying robots.
The real technical challenge here, besides breeding millions of sterilized mosquitos, is actually not the flying robot (drone/UAV) but rather the engineering that needs to go into developing a release mechanism that attaches to the flying robot. In fact, we’re more interested in developing a release mechanism that will work with any number of flying robots, rather than having a mechanism work with one and only one drone/UAV. Aerial robotics is evolving quickly and it is inevitable that drones/UAVs available in 6-12 months will have greater range and payload capacity than today. So we don’t want to lock our release mechanism into a platform that may be obsolete by the end of the year. So for now we’re just using a DJI Matrice M600 Pro so we can focus on engineering the release mechanism.
Developing this release mechanism is anything but trivial. Ironically, mosquitos are particularly fragile. So if they get damaged while being released, game over. What’s more, in order to pack one million mosquitos (about 2.5kg in weight) into a particularly confined space, they need to be chilled or else they’ll get into a brawl and damage each other, i.e., game over. (Recall the last time you were stuck in the middle seat in Economy class on a transcontinental flight). This means that the release mechanism has to include a reliable cooling system. But wait, there’s more. We also need to control the rate of release, i.e., to control how many thousands mosquitos are released per unit of space and time in order to drop said mosquitos in a targeted and homogenous manner. Adding to the challenge is the fact that mosquitos need time to unfreeze during free fall so they can fly away and do their thing, i.e., before they hit the ground or else, game over.
We’ve already started testing our early prototype using “mosquito substitutes” like cumin and anise as the latter came recommended by mosquito experts. Next month, we’ll be at the FAO/IAEA Pest Control Lab in Vienna to test the release mechanism indoors using dead and live mosquitos. We’ll then have 3 months to develop a second version of the prototype before heading to Latin America to field test the release mechanism with our Peru Flying Labs. One of these tests will involve the the integration of the flying robot and the release mechanism in terms of both hardware and software. In other words, we’ll be testing the integrated system over different types of terrain and weather conditions in Peru specifically.
For now, though, our WeRobotics Engineering Team (below) is busy developing the prototype out of our Zurich office. So if you happen to be passing through, definitely let us know, we’d love to show you the latest and give you a demo. We’ll also be reaching out the Technical University of Peru who are members of our Peru Flying Labs to engage with their engineers as we get closer to the field tests in country.
As an aside, our USAID colleagues recently encouraged us to consider an entirely separate, follow up project totally independently of IAEA whereby we’d be giving rides to Wolbachia treated mosquitos. Wolbachia is the name of bacteria that is used to infect male mosquitos so they can’t reproduce. IAEA does not focus on Wolbachia at all, but other USAID grantees do. Point being, the release mechanism could have multiple applications. For example, instead of releasing mosquitos, the mechanism could scatter seeds. Sound far-fetched? 