Archive 03.07.2026

Page 8 of 8
1 6 7 8

#RoboCup2026 – humanoid league day 2

The second day’s play at RoboCup 2026 has drawn to a close with another bumper set of matches. Teams have come from far and wide to take part in the humanoid soccer competition this year, with 17 different countries represented. China is the most represented country, boasting 15 teams across the three divisions. Other countries taking part are geographically widespread, ranging from Colombia to Malaysia, from Germany to Australia.

In advance of the competition, all applying teams provided a video, team description paper, and information about the robots and software that they use. You can see the complete set of these here. As a taster, here is the qualification video from team CAU Mountain&Sea (from China Agricultural University) who are currently leading the small division competition, and are in fourth place in the large division.

Yesterday’s play saw the first issuing of a red card. The robot received two yellow cards for unsafe challenges and was removed from play on safety grounds. You can see a clip of the second foul in question here.

In terms of the vital statistics of the robots, the heaviest and tallest robot is HERoEHS’s ALICE 4 robot, weighing in at 48kg and measuring 160cm tall. There is a big jump down in weight to the second heaviest bot, which is team BigHeroX’s Z4 bipedal humanoid, at 37kg. A number of teams are using the 35kg Unitree G1. At the other end of the spectrum, the lightest and shortest robot is ITAndroids’ Chape, at just 3.8kg and 53cm tall.

With the seeding rounds almost complete (the knockout rounds start tomorrow) the competitions are hotting up nicely. After four competed rounds in the small division, CAU Mountain&Sea is the only team to win all four matches, and sits atop the table with 12 points, and impressively only conceding one goal. Hamburg Bit-Bots and GeoHBots are in second and third place respectively, having both won three games and lost one.

The middle division has fast established itself as one of the most compelling competitions. B-Human has taken a commanding lead, winning all four matches so far with a +35 goal difference. Behind them, there are four teams on nine points: HTWK Robots, Rhoban, whIRLwind Amsterdam, and THMOS.

Over in the large division, teams have completed three rounds of seeding, with just one further round to come tomorrow. At this stage Tsinghua Hephaestus is the only perfect team, sitting on nine points. Behind them, two teams have each won two matches and drawn the third: PCMS-HRG and Robo-Erectus.

Once the seeding rounds have been completed, 12 teams from each division will make it forward to the knockout stages, with the quarter finals taking place in the evening.


Many thanks to JT Genter for providing the photos, videos, and facts for this article.

Robots can now ‘see’ touch thanks to a new color-changing tactile sensor

Engineers at Queen Mary University of London have built a new color-changing tactile sensor, which allows robots to "see" and touch in real-time. The novel idea was invented by Giacomo Sasso, a postdoctoral researcher at the School of Engineering and Materials Science at Queen Mary University of London, and it works by transforming invisible forces into dynamic color patterns. This enables high-resolution maps of contact, strain and pressure to emerge instantly.

Giving drones a sense of ‘pain’ could help them predict instability before it happens

Imagine you're running and you sprain your ankle. The pain makes you gingerly limp the rest of the way home. This is a great example of how nature adapts to failures in a system. The pain tells you: "If you continue running like normal, the injury will only get worse." So you naturally adjust the way you run. Drones currently cannot do this with a worn-out propeller.

Reflections from ICRA 2026

From the 1st-5th June, the robots descended on Vienna. The 2026 IEEE International Conference on Robotics & Automation (ICRA) brought together the top minds in robotics for one short week to showcase the latest technologies, form new collaborations, and exchange ideas. Held at the Messe Wien, a stone’s throw from the bank of the Danube, ICRA proved to be equal parts technological marvel and thought-provoking discussion. 


The host venue for ICRA 2026: Messe Wien, also known as VIECON.

Workshop on robot ethics

My week at ICRA began with the 2nd ICRA 2026 Workshop on Robot Ethics: Ethical, Legal and User Perspectives in Robotics & Automation (WOROBET). WOROBET provided a space for researchers to share ideas, thoughts, and concerns on the future of robot-human interaction, and how to create ethical frameworks to navigate this rapidly changing technology. 

Yasuhisa Hirata, Professor at Tohoku University, began by presenting his vision of a world with physically assistive robots, such as detachable exoskeletons or cycling wheelchairs. These tools can affect people’s sense of self-efficacy and motivation, and there are a host of ethical implications that come with this – how do you help people just enough to build their confidence, without slipping into deception? 

We then heard from Prof. Minoru Asada from Osaka University, who discussed his aim to implement pain signals into robots, so they can experience the world as we do. This was a highly interesting and niche proposition that brought up more ethical questions than we have answers for at the moment. Perhaps most pertinent to a technical conference: is a sense of embodied morality necessary for true intelligence? 

Alan Winfield, Professor of Robot Ethics at UWE Bristol, presented a vision of robotics that acted as a counterweight to Prof. Asada’s: robots as tools, not potential beings. This also somewhat reflects differing attitudes in Eastern vs Western cultures. Thinking about the practical issues we are likely to face in the near term, he outlined a framework for social robot accident investigation. In his view, robot ethics is not just an engineering problem, but requires appropriate social and governance frameworks, just as we do for aviation. His talk also emphasised the risk that programming ethics into robots runs the risk of removing moral responsibility from the roboticist, and can always give rise to unethical robots via malicious hacking. The theme that we must focus on human morality, as opposed to machine morality, was repeated throughout the day. 

After a morning of differing ideas and visions of what robot-human interaction could be, we were invited to ground this into a real robot social care scenario by Praminda Caleb-Solly, Professor of Embodied Intelligence at the University of Nottingham. In this red-teaming exercise, we examined the safety risks and possible mitigations of an assistive robot for a schoolteacher recovering from a stroke at home. Our group discussion circled around human agency: how ethical is it to make design choices for people that take away some of their autonomy, in the name of their best interests? As robots in social care will become a more urgent need in the years to come, these questions may become more salient. 

I left WOROBET with plenty to think about, and a renewed sense of appreciation for the ethicists who are already grappling with the problems that are to come. I hope that progress in robot ethics keeps pace with progress in robotics, so that we are well prepared as robots become more of a part of our daily lives.

Welcome to the jungle – the robot exhibition floor

Every time you entered the exhibition hall, you were greeted by one of the child-sized Booster robots, either playing football, dancing, or demonstrating some kung fu. They were always an endearing welcome to the sea of robots. 

The robots ranged from the endearing to the uncanny, but the common thread was their technical capabilities were astounding. Veteran attendees consistently remarked on how much the robots had improved year on year.

The cute.

The slightly uncanny.

The below clips show the robots that most caught my eye. After admiring a phosphorescent, Stranger Things-esque robotic flower display, I was blown away by the D1-modular robot from Direct Drive. Unlike most robot dogs, it can split into two halves, with the ability to jump, twist, and traverse difficult terrain. 

Sharpa’s North was always a friendly face, waving and making love-hearts at visitors to its booth. Around the back of the booth, you could even challenge it to a round at blackjack. We saw humanoids zipping up rucksacks and trying to fold laundry. Enchanted Tools’ social care robot, Mirokaï, was an unusual sight among the mass of black and steel, with a bright orange body, feline ears and an orange, furry face. Tesollo’s humanoid spent much of its time at ICRA using its long, wavering arms to pick up fruit and drop it into baskets, with impressive dexterity. Vietnamese company Vinrobotics’ humanoid offering was reminiscent of the Cybermen, with its gently wheezing joints, but the team assured me it was much friendlier. The pint-sized Boosters were almost always playing football, not far from their similarly sized Agibot cousins. 

However, humanoids didn’t steal the show this year, as they have done previously. The big trend this year was robotic hands, and the levels of dexterity were truly impressive. Closing the gap between human and robot abilities here would unlock whole new swathes of tasks to automate – and who wouldn’t want a robot folding their laundry? 

Industrial challenges: solving dexterity 

Tackling the dexterity challenge really defined the industrial talks for me this year. A talk by ARIA’s program director, Prof. Jenny Read, demonstrated how the UK government is already laying the groundwork here.

They outlined their funding proposals as part of their Smarter Robot Bodies program, which is split into two branches: robot locomotion and robot dexterity. The robot locomotion branch aims to enable robots to traverse messy, unpredictable physical environments, while the robot dexterity branch will try to break the bottleneck of adept physical manipulation by robotic hands. With the programme set to launch in early 2027, it’ll be exciting to see what kind of innovations this attracts. 

One standout innovation in the realm of dexterity was TARS, a record-setting newcomer in the Chinese robotics market. Co-founded by Dr Ding Wenchao just 18 months ago, TARS has already raised the most funding in angel and pre-seed rounds of any company in the Chinese embodied intelligence sector, and achieved a Guinness World Record for robotic flexible robotic flexible wiring-harness insertion completed in one hour. 

TARS’ DexHand. Image credits: TARS.

DexHand is a 1:1 model of the human hand, even replicating the 21 degrees of freedom we have in the wrist joint and hand. According to TARS, “it can interpret tactile data to distinguish slipperiness, roughness, and hardness in real time and perform 26 English alphabet hand gestures with high-precision finger control.” I was given the chance to control the hand using my own, and was impressed by how well it emulated my own movements. I was also impressed by their humanoid zipping up a backpack – despite the technical abilities of all the robots on the floor, few were able to perform tasks which required such fine motor skills. 

TARS robot zipping up a rucksack.

On Thursday, Dr Wenchao Ding delivered an industry keynote where he presented TARS’ roadmap, charting the path from academia to industrial deployment. With an impressive academic team behind them, TARS may be one to watch. 

Plenary talks

Aside from the wealth of invention and innovation taking place in the exhibition hall, the breadth and depth of academic research at ICRA was fantastic. The plenaries especially gave an insight into the research trends that are currently defining the field. 

Ken Goldberg delivered an electrifying plenary, titled “A Tale of Two Cultures:  Can Agentic Coding Close the Gap?” In this talk, he called for a step change to close the data gap faced by robot manipulation. With the rise of diffusion models and LLMs, it is clear that big data has solved computer vision and language. He challenged the audience – when will the ChatGPT moment for robotics come? With state spaces larger than 50 dimensions in robotics, there is not enough training data to close this gap. Currently, the data required to train vision-language models is equivalent to 100,000 years of real physical experience. 

According to Goldberg, the 2 dominant cultures in engineering – model free “good old fashioned” engineering (GOFE), and, the currently more popular model based engineering.  GOFE encapsulates rigorous engineering methods that pre-date AI, but may have been slightly forgotten about in the AI wave of recent years. He also highlighted how, in his career, he has always been working to bridge the gap between 2 cultures: from science and art; to robotics and automation

He outlined 4 possible solutions to the data gap:

  1. Simulations – these work incredibly well for locomotion and body control, but less so for manipulation due to the number of forces and instabilities involved.
  2. World models – they do not currently properly capture the physics, and hallucinations can be problematic. 
  3. Human teleoperation – this is currently big business and a good way to obtain high quality data. However, the largest dataset is currently only equivalent to a year’s worth of data. 
  4. Real data from functioning robots – this is less commonly used, but can be powerful. 

Prof Goldberg described how he used the fourth approach in his robotic delivery packing company, Ambi Robotics, for 22 years. Picking up bags is an example of variational automation –  one task is done repeatedly, but with different initial conditions each time. From this rich dataset, they created a generative model to train robots in the best way to pick up bags. Here, they close the gap between model-free and model-based methods – Ambi exploits both to achieve industry-leading results. This plenary served as a call for other researchers to use their own production data to do the same. 

During Thursday’s keynote on Robot Learning, Planning & Foundation Models, Stefanie Tellex from Brown University gave a compelling talk titled “Towards Complex Language in Partially Observed Environments”. While current research is bounded in known, predictable scenarios, using action-based language, this does not reflect what the real world is like, nor how people would naturally communicate with robots. Prof. Tellex described her work creating robots that can understand complex, goal-based commands in only  partially observed, dynamic environments, and outlined the grounded Turing test – a reimagining of the Turing test for embodied AI. 

An example of a robot performing a goal-based task in a dynamic environment. Credits: Tellex et al, 2026.

Both plenaries spoke to current pinch points in robotics: data, reasoning, and operating in complex real-world environments. It’ll be interesting to see what solutions are developed in the coming years. 

Science communications crash course

One of my favourite parts of ICRA was delivering the Science Communications Crash Course. Along with Robohub Executive Trustee Sabine Hauert, IEEE Spectrum Senior Editor Evan Ackerman, and IEEE Spectrum Community Manager Kohava Mendelsohn, we gave our guidance on effective science communication to an audience of 100 academics. It was encouraging to see so many people interested in communicating their research effectively – it is a crucial skill, especially in the era of AI and robotics when mainstream narratives can be hijacked by doom-mongering, hype, and corporate interests. More academics communicating their work clearly and neutrally will go a long way to grounding our societal discussion in technical reality, not sci-fi futures. 


Sabine kicking off the science communications crash course. Image credits: Taraja Arnold


Delivering my part of the course. Image credits: Taraja Arnold

Art and robotics 

The arts and robotics section was rich and interesting. There was a lot to visually take in, with constant background music from a robotic saxophone


Masatoshi Hamanaka’s robotic saxophone. Image credits: ©Denes Erdos – Your Event Photographer

PET – marked by a large “PET ME” sign – was a white and orange mass of connecting, rotating pyramids, that gently pulsed and hummed in response to touch, responding by curling towards or away from you depending on how you touched it, The effect was strangely lifelike. 

Rhombus Research presented “Reptile: A Bio-Mimetic Choreography Engine for V2X and A2X Swarms”. This artistic simulation visualises the contracts negotiated within autonomous vehicle and swarm fleets in dreamy blue, browser-based visualisation. Performance artist and former Cirque de Soleil acrobat Silke Grabinger explored human-robot interaction via her piece, AREYOUARE

Silke Grabinger performing AREYOUARE. Image credits: ©Denes Erdos – Your Event Photographer

I’d also like to acknowledge some of the video creators I met there. YouTubers Back to Engineering and the. Amazing, PhD are making some fantastic videos about physical AI. 

Nothing lasts forever – ending with the robot parade 

ICRA 2026 ended with the robot parade, which attracted quite a crowd. See if you can spot the panda, dragon, and headless humanoid!

Seeing all the robots gathered together was a real spectacle. The technology on display was state-of-the-art, and it only improves year on year. What struck me most was the sense that robotics is moving from proving what is possible to tackling the remaining barriers to real-world deployment. Across the exhibition floor, industry keynotes, and plenary talks, the focus was often on the same challenges: dexterity, data, and operating reliably in complex environments. Solving these bottlenecks will open up new avenues to real-world applications of robots. 

While workshops such as WOROBET highlighted important questions around ethics, agency, and governance, the overwhelming emphasis at ICRA 2026 was on capability. Researchers and companies alike are working to close the gap between what robots can do in carefully controlled demonstrations and what they can do in the messy reality of the world outside the lab. Judging by the pace of progress on display in Vienna, that gap may be narrowing faster than many of us expected. I hope that the kinds of conversations around human-robot interaction and ethics that were commonplace at WOROBET and in the Arts exhibition space will become more mainstream – we may need to face the questions that they raise sooner than we think. 

Note: Where image and video credits are not stated, they belong to Ella Scallan.

#RoboCup2026 – humanoid league day 1

Image credit: RoboCup Federation.

RoboCup 2026 kicked off today in Incheon, South Korea, with the league competitions running until 5 July. It’s an exciting time for RoboCup, as there have been some updates to the leagues and competition format. Most prominently, the soccer leagues will have a primary focus on humanoid robots. In a series of daily updates, we’ll be bringing you the latest results, videos, and news from the humanoid soccer league.

This year, the humanoid league is split into three sizes: large division, middle division, and small division. There are 18 teams participating the small division, 16 in the middle, and an impressive 22 competing in the large division.

The first two days of competition are be devoted to the seeding round. Teams play using the Swiss-system of ranking to decide who gets through to the knockout stages.

Livestreams from the different fields of play can be found here.

At the end of the first day of competition, teams in the small and middle divisions have played two games each, with the large division ending the day part way through the second round. Early leaders in the small division, with the full six points from two games are: GeoHBots, CAU Mountain&Sea, and Hamburg Bit-Bots. There are also three teams in the middle division who have claimed the six point haul: B-Human, RoboRoos, and HTWK Robots.

You can watch a short summary of the first day, including some of the robots in action, from KBS News:

This short video from RoboCup gives a flavour of the day’s happenings, which also featured the opening ceremony.

We will be back tomorrow with further updates on competition results, some highlights from the day, and insights into how the event is progressing.

Useful links

How AI-Powered Natural Language Processing Is Reshaping Healthcare?

How Natural Language Processing is Turning the Healthcare Industry in the USA?

The United States’ healthcare sector is experiencing a revolutionary change, and Natural Language Processing (NLP) is at its center. An Artificial Intelligence (AI) arm, NLP is revolutionizing the way clinicians engage with information, documents, and even individuals. From relieving the pain of manual works to enhancing the accuracy of diagnoses and automating billing, NLP is transforming healthcare to make it faster, smarter, and more human. In this article, we’ll explore how NLP is reshaping the way care is delivered, and why it’s quickly becoming a game-changer for healthcare systems across the country.

Transforming Clinical Language into Data that Matters: NLP Is Reading Between the Lines

Previously, it has been challenging for hospitals to manage their data, including doctor notes, discharge summaries, radiology reports, and even call transcripts without proper technology in place. NLP flips that on its head by converting unstructured text into structured, actionable information. Now, rather than manually reading through thousands of words, AI systems can notify:

  • Missed drug interactions
  • Early warning signs of rare diseases
  • Predict the risk factors from the historic data

Use case: Mayo Clinic used NLP to identify suicide risk in teenagers’ months ahead of any intervention that would traditionally occur.

NLP Is the Secret Weapon Behind Smarter Virtual Health Assistants

“Siri, what’s my diagnosis?

With the AI chatbot and telemedicine, NLP drives virtual care in real time. Imagine giving a medical degree to Alexa, but she’s HIPAA-compliant and educated from millions of de-identified patient histories.

Key functions:

  • Patient symptom understanding
  • Layman-clinical and clinical-layman translation
  • Clinician-automated chart notes generation
  • Triaging and appointment scheduling support

This is not convenience; this is life-saving automation in rural or underprivileged communities where there is a thin pickup of doctors on the ground. 

NLP Is Cracking Down on Fraud and Billing Nightmares

“Killer Robots for Insurance Claims”

The American billing process is famously a tangle of ICD-10 codes and confusing forms. NLP technology is now revolutionizing this space by automating claims generation, detecting fraudulent schemes, and ensuring equitable insurance reimbursements. With the expertise of app development companies, these advanced NLP-driven solutions are being integrated into healthcare systems, making billing faster, more accurate, and far less stressful for both providers and patients.

NLP Is Reading Medical Literature Faster Than Any Human Could

New research studies are being published every 26 seconds. No physician can ever hope to keep up with that rate, but NLP can.

AI models are already reading thousands of daily studies today, mining trends, drug interactions, trial outcomes, and guidelines and delivering live feeds into clinical decision aids. It informs physicians in real time and rescues them from ignorance errors. 

10 Top NLP Trends in Healthcare

  1. Generative AI for Clinical Documentation

Physicians are applying AI-driven instruments such as ambient scribe tech to generate clinical notes automatically from dictations, preventing burnout and enabling them to spend more time with patients.

  1. Unstructured Data Mining

Physicians are investing in NLP to make inferences from unstructured data — like EHR narratives, radiology reports, and pathology reports — to support diagnosis and care planning.

  1. Voice-Activated Assistants

NLP-powered virtual assistants are being trained to assist with real-time engagement of patients and staff, answer questions, handle scheduling, and assist treatment decision-making.

  1. Clinical Decision Support with AI

NLP is being utilized to gather relevant medical history, symptoms, and laboratory results to support physicians with evidence-based, real-time decision-making during patient interactions.

  1. Patient Sentiment and Emotion Analysis

Hospitals are applying NLP to process patient feedback, surveys, and even SMS to identify dissatisfaction, anxiety, or risk, leading to better patient experience and mental healthcare.

  1. Population Health & Social Determinants Analysis

NLP solutions are able to identify concealed social or behavioral health illnesses (e.g., housing instability or substance abuse) in free-text reports to inform public health practitioners to anticipate threats in communities.

  1. Monitoring Bias and Fairness

New NLP models are coming in with bias detection features to treat all on par, regardless of race, gender, or language community, a good step by regulators and stakeholders in health equity.

  1. EHR System Integration

NLP is being increasingly embedded in top Electronic Health Record systems (such as Epic and Cerner) to enable search, workflow automation, and usability of data for clinicians.

  1. Multilingual NLP Models

Multilingual NLP solutions are being utilized in multicultural-population hospitals to enable Spanish, Mandarin, Arabic, and other language-speaking patients, bridging communication care gaps.

  1. Real-Time Clinical Analytics

Real-time NLP dashboards are increasingly being deployed in ICUs and ERs to monitor symptoms, risk, and treatment outcomes to enable teams to respond more quickly during emergencies.

 

These innovations are delivering more care, fewer mistakes, and lower bills — and they illustrate how NLP is emerging as an integral part of intelligent, data-based medicine.

How Much Does NLP Development for Healthcare Cost?

Developing Natural Language Processing (NLP) healthcare solutions in the US is a fairly costly based on numerous things, from the size of a project and the data complexity to requirements like suitability with HIPAA guidelines. A few factors that have an impact of NLP development cost are mentioned below:

Creating an NLP system for American medicine can be expensive, based on what the system has to do. A simple tool, like one that helps doctors write automatically or transcribe, can cost $100,000 to $300,000.

Sophisticated systems that look at medical records or help with clinical decisions can range from $500,000 to millions of dollars.

Compliance with HIPAA is a big reason for the expense. Healthcare data is confidential, and hence any software developed must be subject to very strict regulations to protect patient confidentiality. That costs extra in terms of security, legal effort, and regular system testing.

Cloud computing, software licenses, and supercomputers utilized in training AI models may run into thousands of dollars per month. Once the system is established, it has to be serviced and upgraded from time to time, generally 15–25% of the project cost annually.

Generally, NLP app development companies can cost from $150,000 to $500,000. It can be expensive, but it saves time, decreases medical errors, and enhances patient care in the end.

Conclusion

NLP Isn’t the Future of Healthcare, It’s the Now.

From translating complex EHRs to helping patients schedule appointments, NLP is woven into the healthcare sector. It is cost- and time-efficient and picks up issues around privacy, accuracy, and equity. If you have plans of developing NLP applications for healthcare, connect with us.

Contact us to know more about How AI-Powered Natural Language Processing Is Reshaping Healthcare? Book Executive AI Briefing →

 

[contact-form-7]

A decade of open source at DataRobot: from predictive AI to the agent lifecycle

A decade of open source at DataRobot: from predictive AI to the agent lifecycle

Every era of DataRobot has shipped open source. The latest open-source contributions from DataRobot map directly onto where agents actually break in production.

Three benchmark charts side by side: syftr accuracy vs. latency, Token Pool latency over time, and JointFM Sharpe improvement vs. market volatility

Building an agent has never been easier. Pick a framework, wire up a model and a retriever, add a few tools, and a demo is running by lunch. The trouble starts after the demo. The workflow you guessed at turns out to be neither the most accurate option nor the cheapest one. The agent has to make a judgment call under uncertainty and has no fast way to reason about risk. And the moment more than one team starts using it, the inference bill and the latency both go sideways.

These are not framework problems. They are lifecycle problems, and they surface at three distinct stages: designing the workflow, reasoning under uncertainty at runtime, and serving the result to real users at scale.

None of this is new territory. Open source at DataRobot has never been a side quest. It has tracked the platform’s evolution stage by stage: teaching predictive AI in the open, then giving teams programmatic ownership of AutoML, and now shipping the actual infrastructure for each place agents go to production.

A decade of showing the work

The habit goes back to 2014, when the team open sourced its top-finishing code from the KDD Cup, alongside blog tutorials on gradient boosting, scikit-learn, and regression in statsmodels. The tutorials for data scientists repository, and later a run of generative AI accelerators, grew out of the same instinct: the only way to really understand AI is to build it, so hand people working code instead of a white paper. All of it sat on top of the R and Python SDKs, which is what turned a trial account into something people could script against instead of just click through.

Education answers “how do I learn this.” The next question is “how do I trust what got built,” and the answer was orchestration. The Pulumi provider and the accompanying CLI let a workflow be defined as code and rerun on someone else’s machine with the same result, turning AutoML from a black box into an exportable, auditable record. Blueprint Workshop, a Python client for constructing and editing blueprints programmatically, extended the same idea to the modeling layer itself: preprocessing, algorithms, and post-processing as code, not just as nodes in a UI.

Ownership was the logical next step after orchestration. Custom Models and Custom Tasks, built on the open-source DRUM framework, let teams bring their own pretrained models and preprocessing steps into a deployment and get monitoring, governance, and a leaderboard for free. Composable ML on top of Custom Tasks meant a blueprint could mix the platform’s own algorithms with a team’s proprietary preprocessing, without forcing a choice between the two.

The connective tissue between that era and this one is Pulumi. The same declarative pattern that once documented a predictive pipeline now provisions agent infrastructure: agent templates for CrewAI, LangGraph, and LlamaIndex ship with Pulumi wired in by default. The tools changed. The commitment to a code path instead of a walled garden didn’t.

The agent lifecycle, and where it breaks

It helps to name the stages before naming the tools. An agent moves through a predictable arc. You design the workflow that defines how it retrieves, reasons, and responds. At runtime, it has to reason about an uncertain world well enough to act. And the platform has to serve that agent to many tenants without breaking service level objectives or the budget. Each stage has a hard question attached, and three shipped projects, plus one still in review, exist to answer them.

syftr: design the workflow before you guess

The first decision in any RAG or agentic build is also the one teams skip: which configuration to use. Which synthesizing LLM, which embedding model, which retriever, what chunk size, whether to add reranking, whether the flow should be agentic at all. The space runs past ten to the twenty-third unique configurations, and every choice trades accuracy against latency against cost. Most teams pick a reasonable-looking default and never find out how far it sits from the frontier.

syftr searches that space instead of guessing. It uses multi-objective Bayesian optimization to find Pareto-optimal flows: the configurations where accuracy cannot improve without paying more, and cost cannot drop without losing accuracy. A domain-specific early-stopping mechanism prunes clearly suboptimal candidates before they burn through an evaluation budget, cutting search compute by 60 to 80%. On industry-standard RAG benchmarks, it identifies workflows that cut cost by up to 13 times with only marginal accuracy trade-offs.

syftr doesn’t replace judgment. It gives a data-driven way to navigate a design space too large to reason about by hand, searching across 10 proprietary and open-source LLMs, 13 embedding models, four prompt strategies, three retrievers, and four text splitters, and it produces production-ready pipeline code at the end.

pip install git+https://github.com/datarobot/syftr.git

JointFM: give the agent a quant for runtime decisions

Designing the workflow gets an agent built. It doesn’t help the agent make a hard call at runtime. Rebalancing a portfolio, hedging a supply chain, dispatching energy on a grid: these are decisions under uncertainty that depend on the full joint distribution of many coupled outcomes, including how they move together in the tails. Classical quantitative methods model this well but are slow and brittle. Standard time-series foundation models are fast but forecast each series in isolation, missing the cross-variable dependencies that define systemic risk.

JointFM, a foundation model from DataRobot Research, closes that gap. Pretrained on a stream of synthetic stochastic differential equations rather than fit to one dataset, it learns the underlying physics of stochastic dynamics and predicts the full joint distribution of future outcomes in a single forward pass: 10,000 samples across 10 targets and 63 horizons in roughly 10 milliseconds on a single GPU, with a 21.1% reduction in energy loss against the strongest classical baseline in zero-shot evaluation.

That speed is what makes it relevant to agents. When a decision-making agent can resimulate the future in milliseconds, risk-aware reasoning becomes something it does inline, on every decision, instead of a nightly batch job a person waits on. An initial finance application matched the risk-adjusted returns of a classical benchmark while replacing an overnight process with a real-time one. The model is domain-agnostic by construction, so the same capability extends to energy and logistics.

Token Pool: serve every tenant without starving the ones that matter

A well-designed agent with sharp runtime reasoning still has to run somewhere, usually alongside everyone else’s. Multi-tenant inference hits a wall here. Dedicated endpoints strand GPU capacity on idle models. Rate limits treat every token as equal, even though one request can cost an order of magnitude more GPU time than another. Neither approach lets idle capacity be borrowed, and both fall apart under the bursts that characterize real inference traffic. The familiar result: one team’s batch job floods the endpoint, and everyone’s production latency spikes.

Token Pool fixes this at the API gateway, without touching the inference runtime underneath. It expresses capacity in inference-native units, token throughput, KV cache, and concurrency, rather than machine or pod counts. Tenants hold entitlements to a share of a pool, and service classes (dedicated, guaranteed, elastic, spot, and preemptible) set the protection ordering during contention. A debt-based fairness mechanism gives temporarily throttled workloads compensatory priority later, so no tenant is starved and none monopolizes the pool. It runs as a Kubernetes-native layer above vLLM or TensorRT-LLM.

In overload testing, Token Pool held sub-1.2 second P99 time-to-first-token for guaranteed workloads by selectively throttling spot traffic, while a baseline with no admission control degraded past 19 seconds across every workload. For anyone responsible for consumption-based economics or API governance, this is the missing primitive: capacity expressed in units that match what inference actually costs.

kubectl apply -f examples/sample-tokenpool.yaml
kubectl apply -f examples/sample-entitlement.yaml

What’s next: closing the loop

The three shipped projects operate as separate links today. Design-time search runs once. Runtime reasoning runs blind to how the serving layer is performing. The serving layer enforces policy without feeding anything back upstream. The workflow syftr found last quarter isn’t necessarily optimal against this month’s traffic, models, and prices.

The next open-source project connects production telemetry, the real cost, latency, and quality signals coming off the serving layer, back to the optimization layer, so workflows get re-evaluated against production reality instead of a single offline benchmark. It’s still in review, so it isn’t named yet, but it’s the natural fourth stage after design, reason, and serve.

Get started

  • Build: install syftr with pip install git+https://github.com/datarobot/syftr.git and run the starter search
  • Build: stand up Token Pool against a local Kind cluster, no GPU required
  • Read: the instant portfolio optimization walkthrough for JointFM

A hands-on guide for each follows next in this series: running a first syftr search and reading the Pareto frontier, calling JointFM for real-time forecasts inside an agent, and standing up Token Pool to protect a production workload from a noisy neighbor. Start with whichever stage of the lifecycle is hurting most.

The post A decade of open source at DataRobot: from predictive AI to the agent lifecycle appeared first on DataRobot.

2026 BAIR Graduate Showcase

Congratulations to the Berkeley Artificial Intelligence Research (BAIR) Lab class of 2026! This year, BAIR celebrates another remarkable group of Ph.D. graduates whose curiosity, creativity, and perseverance have pushed the frontiers of artificial intelligence and machine learning.

Their work spans the breadth of modern AI — robotics and embodied intelligence, large language models and reasoning, computer vision, generative modeling, AI safety, human-AI interaction, AI for science and healthcare, and much more. Along the way, they have published influential research, built systems with real-world impact, mentored their peers, and shaped the BAIR community for the better.

Now they are headed everywhere ideas travel: to faculty and postdoctoral positions, to industry research labs, and to startups of their own founding — and several are still exploring what comes next and would love to hear from you.

Please join us in celebrating the achievements of these wonderful graduates. We are proud of everything they have accomplished at Berkeley, and we can’t wait to see what they do next!

Read More
Page 8 of 8
1 6 7 8