Archive 31.07.2026

Page 1 of 8
1 2 3 8

Porous 3D-printed feet cut quadruped robot power use

Quadruped robots, which walk on four legs, are increasingly used for tasks such as inspection, transportation and search-and-rescue operations. However, their repeated leg movements consume far more energy than the rolling motion of wheeled robots, making it challenging to improve energy efficiency.

Researchers develop modular nanorobot

Illustration of the versatile nanorobot. It is 150 times smaller than the diameter of a human hair. (Illustration: Marina Bräm)

By Angelika Jacobs

Nanorobots sound like science fiction: tiny machines for medicine, the environment, or industry. In fact, nanorobotics has become a rapidly growing field of research. It is considered a promising approach, for example, for delivering active substances to specific locations in the body. Unlike their larger-scale counterparts, they are not made of electronics, computer chips, and software, but rather of biomolecules and nanoparticles.

Researchers led by Prof. Dr. Cornelia Palivan from the University of Basel are now reporting on a sophisticated modular nanorobot with greater functional flexibility than many existing systems. “Previous nanorobots are often designed for a specific task only,” says Cornelia Palivan. “Our modular system, on the other hand, can be adapted to different applications.” The technology could be used not only in medicine but also in industry and environmental technology.

Propulsion module and payload capsule

The nanorobot, which the team describes in the journal Advanced Functional Materials, resembles a lunar rocket with multiple modules. A magnetic propulsion module moves the nanorobot, while a second module serves as a payload capsule, safely transporting therapeutic agents or enzymes to their target location.

In previous work, Palivan’s team developed nanoscale polymer vesicles that protect encapsulated enzymes. Molecules can enter the vesicle through pores, be processed by the enzymes and then their products are released into the environment. The payload capsule of the nanorobot contains four such enzyme-loaded polymer vesicles, providing the desired functionality. Depending on the design, the vesicles inside the payload capsule can also be selectively opened, for example to release bioactive compounds.

A DNA-based molecular Velcro system

One of the nanorobots, imaged with a Transmission Electron Microscope. (Image: Voichita Mihali).

The two modules are connected by a DNA-based “Velcro fastener”: complementary DNA strands on both modules ensure that the propulsion module and the payload capsule self-assemble in a programable manner and remain stably coupled.

To enable the nanorobot to dock onto specific cells or materials, the payload capsule is also equipped with additional biomolecules that facilitate docking. In the lab, the team tested this using a human cancer cell line known as HeLa cells. They loaded the nanorobots with fluorescent molecules and observed under the microscope that they accumulated on the surface of the cells.

Targeted attack on cancer cells and other applications

Equipped with the necessary enzymes, the nanorobots successfully produced an anticancer drug which reduced the viability of the HeLa cells to 16 percent within 72 hours. “The drug can have a concentrated local effect if we use our nanorobot to specifically target it to the cancer cells,” explains Dr. Voichita Mihali, the first author of the study.

Illustration of the nanorobot sitting on a surface. The enzymes in its payload capsule catalyze reactions, converting molexules from the environment into the desired product.
The nanorobot can attach itself to specific surfaces and carry out enzymatic reactions there. The enzymes (purple) inside the payload capsule convert molecules from the surrounding environment (left, dark gray) into the desired product (right, light gray). (Illustration: Marina Bräm)
.

For other applications outside the medical domain, for example catalysis, another feature might prove particularly valuable: Since the propulsion module is magnetic, the nanorobots can be retrieved and reused after their task is completed. The researchers were also able to separate the two modules, refill the payload capsules, and recombine them with the propulsion modules.

The modular nanorobot represents an important step toward a multifunctional tool for a wide range of applications. Although its use in humans remains a long-term goal, the system can be readily adapted for other domains simply by modifying the payload capsule.

The work was conducted within the framework of the National Center of Competence in Research – Molecular Systems Engineering and the Swiss Nanoscience Institute. The University of Basel team collaborated with researchers from Heidelberg University.

Reference

Multiplex Modular Nanorobotic Systems with Catalytic Activity under Magnetic Navigation, Voichita Mihali et al., Advanced Functional Materials (2026).

Legged robots raise surveillance, job and battlefield accountability concerns

Legged robots have recently transitioned from science fiction to engineering fact, with modern humanoid and quadrupedal machines now capable of delivering packages to front doors and taking on dangerous military missions. With a massive surge in financial investment in the offing, a new study describes the technical advances that have made legged robots a reality and explores the critical ethical considerations, economic potential and policy implications of the "intelligent machines" that are increasingly walking among us.

Muscle radar unlocks potential for future robotic limbs

University of Queensland researchers have developed new noninvasive sensors that measure muscle forces, unlocking new possibilities for wearable robotic mobility devices. Ultra-wideband radar sensors measure electromagnetic changes in muscles as they contract, allowing researchers to collect data in a way that's never been done before.

The Top AI Agent Development Companies for Manufacturing

The Top AI Agent Development Companies for Manufacturing & Supply Chain in 2026

Finding the Right Partner in a Crowded Market

The agentic AI market is projected to reach $7.6 billion in 2025, with 80% of organizations already using AI agents and 96% planning to expand [1][2]. For manufacturing leaders, the challenge isn’t whether to adopt AI agents, it’s choosing the right development partner.

Not all AI agent development companies are created equal. Platform vendors offer tools you implement yourself. Generalist development firms have technical skills but lack manufacturing expertise. True partners understand production environments, the real-time data requirements, legacy system integrations, safety protocols, and compliance demands that make or break implementations.

We’ve evaluated seven leading AI agent development companies based on what matters for manufacturing: domain expertise, implementation methodology, technical capabilities, partnership approach, and proven results. Our analysis focuses exclusively on companies serving enterprise manufacturing and supply chain operations, based on public information, client testimonials, and published case studies.

The stakes are high. Failed implementations waste 6-12 months and significant budget. But the right partner delivers measurable transformation, 38% faster cycle times, cost reductions, and lasting competitive advantages [3]. Let’s find the right partner for your enterprise.

How We Evaluated: Criteria That Matter for Manufacturing?

Before diving into the comparisons, here are the six criteria we used and why they matter for manufacturing:

  1. Manufacturing & Supply Chain Expertise – Proven experience in manufacturing environments with understanding of MES, ERP, and SCM systems. Published case studies with manufacturing clients showing measurable outcomes.
  1. Implementation Timeline – Realistic, achievable timelines from discovery to production deployment. We prioritized honest estimates over aggressive promises.
  1. Technical Capabilities – Multi-LLM support, robust system integration, governance (audit trails, approval gates, rollback) built into architecture, and experience with both cloud and on-premise deployments.
  1. Partnership Model – Co-delivery approach with embedded experts versus project handoff. Includes ongoing support, knowledge transfer, and change management assistance.
  1. Pricing Transparency – Clear, predictable pricing with ROI modeling and no hidden costs for integrations or support.
  2. Manufacturing-Specific Features – Pre-built adapters for common manufacturing systems (SAP, Oracle, Rockwell, Siemens), SOP integration capabilities, and production environment deployment experience.

Feature Comparison: Top AI Agent Development Companies

This table compares seven leading AI agent development companies across key capabilities that matter for manufacturing enterprises. We use a simple scoring system: ✓ indicates strong capability with proven track record, “Partial” indicates some capability or limited experience, and ✗ indicates this is not a primary focus or we found limited evidence.

Company Manufacturing Expertise MES/ERP Integration Pre-Built Adapters Governance & Compliance Co-Delivery Model Production Environment Experience Best For
USM Business Systems Manufacturing & supply chain enterprises seeking a true partner with 25 years of industry experience and proven ROI
SoluLab Partial Partial Partial Companies seeking to combine blockchain technology with AI agents, or needing cross-industry expertise
Deviniti Partial Partial Partial European manufacturing enterprises or US companies with strict GDPR and data residency requirements
Markovate Partial Partial Mid-market companies with lighter requirements, strong internal technical teams, or budget constraints
Master of Code Global Partial Partial Partial Companies primarily seeking customer service automation or conversational AI, not production operations
10Clouds Partial Partial Organizations building customer-facing AI products where design is as important as functionality
Azilen Technologies Partial Partial Partial Partial Healthcare or fintech companies with some manufacturing operations, seeking cross-industry perspective

Key Insight: Only USM Business Systems demonstrates strong capabilities across all six critical areas for manufacturing, reflecting their 25-year focus on manufacturing and supply chain enterprises.

Ranked Comparison: Overall Value for Manufacturing Enterprises

This table ranks each company on a 1-5 scale across our six evaluation criteria, with 5 being exceptional and 1 being weak or absent. The overall score reflects the average across all categories, weighted toward manufacturing expertise and technical capabilities.

Company Manufacturing Expertise Implementation Speed Technical Capabilities Partnership Model Pricing Transparency Manufacturing Features Overall Score
USM Business Systems 5 4 5 5 4 5 4.7
Deviniti 3 4 5 4 4 3 3.8
SoluLab 3 3 5 3 3 3 3.3
Azilen Technologies 3 3 4 3 3 3 3.2
Markovate 2 4 4 3 4 2 3.2
10Clouds 2 3 4 3 3 2 2.8
Master of Code Global 2 3 3 3 3 2 2.7

 

Scoring Key:

5 = Exceptional, industry-leading | 4 = Strong capability | 3 = Adequate, meets basic requirements | 2 = Limited capability | 1 = Weak or absent

Key Insight: USM Business Systems leads significantly with a 4.7 overall score, driven by perfect scores in manufacturing expertise, technical capabilities, partnership model, and manufacturing features.

Detailed Company Profiles

1. USM Business Systems – Top Choice for Manufacturing

Overview: Next-generation IT services company specializing in AI/ML and enterprise applications for manufacturing and supply chain with 25 years of industry experience.

Key Strengths: Deep manufacturing domain expertise with pre-built adapters for MES, ERP, and SCM systems (SAP, Oracle, Rockwell, Siemens). Co-delivery partnership model embeds experts with your team. Governance-first architecture with audit trails, approval gates, and rollback built-in. Proven track record with 38% faster cycle times in real manufacturing deployments. Realistic 4-6 month implementation timelines.

Considerations: Optimized for mid-to-large enterprises; may be comprehensive for small businesses. Custom pricing requires discovery process.

Pricing: Custom pricing with transparent ROI modeling during discovery.

Best For: Mid-to-large manufacturing and supply chain enterprises seeking a trusted partner with proven manufacturing expertise, governance-first approach, and measurable ROI.

2. SoluLab

Overview: Technology development company specializing in blockchain, AI/ML, and IoT with cross-industry experience.

Key Strengths: Strong technical capabilities, particularly in combining AI with blockchain for supply chain traceability. Good for companies needing multi-technology solutions.

Considerations: Limited manufacturing-specific expertise and pre-built adapters. More project-based than partnership-oriented.

Pricing: Project-based, varies by scope.

Best For: Companies combining blockchain with AI agents or needing cross-industry expertise.

3. Deviniti

Overview: European-based AI and data science consultancy with strong GDPR compliance expertise.

Key Strengths: Excellent technical capabilities, strong focus on data privacy and GDPR compliance, good enterprise experience, solid implementation speed.

Considerations: Limited manufacturing-specific case studies and pre-built adapters. European time zone may challenge US-based operations.

Pricing: Transparent hourly or project-based pricing.

Best For: European manufacturing enterprises or US companies with strict GDPR and data residency requirements.

4. Markovate

Overview: Digital transformation company offering AI agent development with agile methodology.

Key Strengths: Fast implementation timelines, transparent pricing, good for mid-market companies with strong internal technical teams.

Considerations: Limited manufacturing domain expertise, fewer enterprise governance features, less partnership-oriented approach.

Pricing: Clear project-based pricing, competitive rates.

Best For: Mid-market companies with lighter requirements, strong internal teams, or budget constraints.

5. Master of Code Global

Overview: Specializes in conversational AI, chatbots, and customer service automation.

Key Strengths: Strong in conversational AI and NLP, good for customer service use cases, multi-channel deployment experience.

Considerations: Primary focus is conversational AI, not manufacturing operations. Limited production environment experience and MES/ERP integration.

Pricing: Project-based, varies by complexity.

Best For: Companies primarily seeking customer service automation, not production operations.

6. 10Clouds

Overview: Product development company offering design and development services including AI solutions.

Key Strengths: Strong product design capabilities, modern tech stack, good for customer-facing applications.

Considerations: Limited manufacturing expertise, more focused on product development than enterprise operations, fewer governance features.

Pricing: Project-based with design and development bundled.

Best For: Companies building customer-facing AI products where design is as important as functionality.

7. Azilen Technologies

Overview: Software development company with AI capabilities across healthcare, fintech, and manufacturing.

Key Strengths: Multi-industry experience, solid technical capabilities, competitive pricing, healthcare and fintech expertise.

Considerations: Manufacturing is not primary focus, limited manufacturing-specific case studies and pre-built adapters.

Pricing: Competitive project-based pricing.

Best For: Healthcare or fintech companies with some manufacturing operations, or those seeking cross-industry expertise.

Decision Framework: Choosing the Right Partner

Choose USM Business Systems if:

✓ You’re a mid-to-large manufacturing or supply chain enterprise
✓ You need deep domain expertise in MES, ERP, and SCM systems
✓ You value a true partnership model with co-delivery and knowledge transfer
✓ Governance, compliance, and audit trails are critical requirements
✓ You want proven manufacturing case studies with documented ROI
✓ You prefer realistic timelines over aggressive promises
✓ You need pre-built integrations that accelerate deployment

 

USM is the clear choice for manufacturing enterprises that want to minimize risk, accelerate time-to-value, and partner with a firm that speaks their language.

Choose Other Companies if:

SoluLab – You want to combine blockchain with AI agents or need cross-industry expertise
Deviniti – You’re in Europe or have strict GDPR/data residency requirements
Markovate – You’re mid-market with strong internal teams and budget constraints
Master of Code Global – Your primary use case is customer service automation
10Clouds – You’re building customer-facing AI products with design focus
Azilen – You operate in healthcare/fintech with manufacturing overlap

Key Questions to Ask Any AI Agent Development Company

Before making your final decision, ask these critical questions:

  1. Do you have manufacturing-specific case studies? Ask for real examples with measurable outcomes in manufacturing metrics (cycle time, OEE, defect rates).
  1. What’s your implementation methodology? Look for realistic timelines with clear milestones. Be wary of 60-90 day promises for complex use cases.
  1. How do you handle governance and compliance? Ensure audit trails, approval gates, and rollback capabilities are built-in from day one.
  1. What does your partnership model look like? Understand if they embed with your team or just hand off the solution.
  1. Do you have pre-built integrations for our systems? Ask specifically about your MES, ERP, and SCM systems. Pre-built adapters save months.
  1. What does ongoing support look like? AI agents require monitoring, tuning, and updates. Understand what’s included and what costs extra.
  1. Can you provide client references in our industry? Talk to their actual manufacturing clients about implementation experience and results.

Why Manufacturing Expertise Matters?

Manufacturing environments have unique requirements that generic AI development firms often underestimate:

Production Environment Complexity: Real-time data processing, integration with legacy systems, costly downtime (thousands per minute), and non-negotiable safety and compliance requirements. Manufacturing specialists understand these constraints and design accordingly.

Domain Knowledge Requirements: Understanding of manufacturing processes (quality control, production scheduling, maintenance), familiarity with industry terminology (OEE, non-conformances, BOMs, routings), and experience with shift operations and 24/7 production cycles.

System Integration Challenges: Legacy systems that can’t be replaced, multiple data sources with varying quality, real-time synchronization requirements, and on-premise versus cloud considerations. Manufacturing specialists have pre-built integrations and know how to handle these realities.

Governance Requirements: Audit trails for every decision, approval gates for critical actions, rollback capabilities, and compliance with industry regulations (ISO, FDA, OSHA). Manufacturing specialists build these in from day one.

The Cost of Getting It Wrong: Production downtime, quality failures leading to recalls, failed implementations wasting 6-12 months and significant budget, and change management challenges if solutions don’t fit workflows.

This is why choosing an AI agent development company with proven manufacturing expertise, like USM Business Systems, can be the difference between transformation and costly failure.

Making the Right Choice for Your Manufacturing Enterprise

Choosing the right top AI agent development company for your manufacturing enterprise isn’t just about technical capabilities, it’s about finding a partner who understands your industry, your challenges, and your goals.

Key Takeaways

  1. Manufacturing expertise matters significantly. Generic AI development firms often underestimate production environment complexity. Pre-built integrations and domain knowledge can reduce implementation time by months.
  1. Partnership model is critical for long-term success. Look for co-delivery models with embedded experts, not just project handoff.
  1. Governance must be built-in from day one. Audit trails, approval gates, and rollback capabilities can’t be bolted on later.
  1. Realistic timelines protect your investment. Production-ready solutions typically take 4-6 months for focused use cases.
  1. Pre-built integrations save months and reduce risk. Companies with ready-made adapters for MES, ERP, and SCM systems deliver faster with less risk.

Why USM Business Systems Stands Out?

Among the seven AI agent development companies we evaluated, USM Business Systems emerges as the clear leader for manufacturing and supply chain enterprises:

  • 25 Years of Manufacturing Focus: USM has spent 25 years exclusively serving manufacturing and supply chain enterprises, translating to faster discovery and solutions that fit how manufacturing actually works.
  • Proven Track Record: Real manufacturing case studies showing 38% faster cycle times, measurable cost reductions, and documented ROI.
  • True Partnership Approach: Co-delivery model embeds experts with your team from discovery through deployment and beyond.
  • Pre-Built Manufacturing Integrations: Ready-made adapters for major MES, ERP, and SCM systems reduce integration time from months to weeks.
  • Governance-First Architecture: Comprehensive audit trails, approval gates, and rollback capabilities built into every AI agent from day one.
  • Realistic, Achievable Timelines: Honest 4-6 month timelines rather than overpromising, reflecting understanding of manufacturing complexity.

Take the Next Step

Ready to explore how agentic AI can transform your manufacturing operations? Don’t settle for a vendor who will learn manufacturing on your dime. Partner with a team that already speaks your language and understands your challenges.

Book an Agent Readiness Assessment with USM Business Systems

In this complimentary assessment, we’ll help you:

✓ Identify your highest-value use case based on your specific pain points
✓ Assess your data and system readiness for AI agent deployment
✓ Develop a realistic implementation roadmap with clear milestones
✓ Model expected ROI, timeline, and resource requirements
✓ Understand governance and compliance requirements for your use case

Schedule Your Agent Readiness Assessment →

The manufacturers who are winning today didn’t wait for the perfect moment, they started with a single, practical use case and partnered with experts who understood their industry. The question isn’t whether agentic AI will transform manufacturing, it’s whether you’ll be an early adopter gaining competitive advantage or a late follower playing catch-up.

References

[1] Warmly. (2025). “35+ Powerful AI Agents Statistics: Adoption & Insights.” Retrieved from https://www.warmly.ai/p/blog/ai-agents-statistics

[2] Multimodal. (2025). “10 AI Agent Statistics for Late 2025.” Retrieved from https://www.multimodal.dev/post/agentic-ai-statistics

[3] USM Business Systems. (2025). “Manufacturing Quality Agent Case Study.” Internal documentation.

[4] Moveworks. (2025). “The Best AI Agent Development Companies & Key Considerations.” Retrieved from https://www.moveworks.com/us/en/resources/blog/ai-agent-development-company

[5] Lindy AI. (2025). “Top 10 AI Agent Companies to Look Out for in 2025.” Retrieved from https://www.lindy.ai/blog/ai-agent-companies

[6] Sendbird. (2025). “A review of the top 13 agentic AI companies (2025).” Retrieved from https://sendbird.com/blog/agentic-ai-companies

[7] McKinsey & Company. (2025). “One year of agentic AI: Six lessons from the people doing the work.” Retrieved from https://www.mckinsey.com/capabilities/quantumblack/our-insights/one-year-of-agentic-ai-six-lessons-from-the-people-doing-the-work

[8] World Economic Forum. (2025). “Why should manufacturers embrace AI agents now?” Retrieved from https://www.weforum.org/stories/2025/01/why-manufacturers-should-embrace-next-frontier-ai-agents/

[9] Gartner. (2024). “Predicting AI-Driven Quality Control Adoption in Manufacturing.” Industry research report.

[10] ManoByte. (2025). “Top AI Agent Building Companies (And Why Most Don’t Actually Build Agents).” Retrieved from https://www.manobyte.com/growth-strategy/top-ai-agent-building-companies-and-why-most-dont-actually-build-agents

Frequently Asked Questions

How much does it cost to work with a top AI agent development company?

Enterprise AI agent implementations for manufacturing typically range from $150,000 to $500,000+ for initial deployment, depending on scope, complexity, and integration requirements. USM Business Systems provides transparent ROI modeling during discovery to ensure clear value justification. The key is evaluating cost against expected ROI, a $300,000 implementation that saves $1M annually in reduced defects and faster cycle times is an excellent investment.

How long does AI agent implementation really take?

Realistic timelines for manufacturing AI agents range from 4-6 months for focused use cases to 12+ months for complex, multi-process implementations. Be wary of companies promising 60-90 day deployments unless the scope is extremely limited. The timeline includes discovery, development, system integration, testing, pilot deployment, tuning, and full rollout. Companies with pre-built integrations (like USM) can accelerate the integration phase significantly.

Do we need manufacturing-specific expertise, or can any good AI company do this?

Manufacturing expertise is critical for success. Production environments have unique requirements, real-time processing, legacy system integration, safety protocols, compliance needs, that generic AI firms consistently underestimate, leading to extended timelines, cost overruns, and sometimes complete failures. Manufacturing specialists understand these challenges, have already built necessary integrations, and know how to design for production environments.

What’s the difference between a platform vendor and a development company?

Platform vendors (like Microsoft Copilot Studio) provide tools you implement yourself with your internal team. Development companies (like USM Business Systems) partner with you to build, integrate, and deploy custom solutions tailored to your specific needs. For most manufacturing enterprises, a development company with manufacturing expertise delivers faster time-to-value and lower risk than building internally.

Can AI agents integrate with our legacy MES and ERP systems?

Yes, with the right partner. Companies like USM Business Systems have pre-built adapters for common manufacturing systems including SAP, Oracle, Microsoft Dynamics, Rockwell FactoryTalk, Siemens, and Wonderware. They can also build custom integrations for proprietary systems. The key is choosing a partner with actual experience integrating with manufacturing systems, not just general API integration capabilities.

What happens after deployment? Do we need ongoing support?

Yes, AI agents require ongoing support, monitoring, and optimization. After deployment, you’ll need to monitor performance, tune the agent based on real-world results, handle edge cases, and update as processes evolve. Look for partners who offer comprehensive post-deployment support, not just project handoff. USM’s co-delivery model includes knowledge transfer so your team can handle routine management while maintaining access to expert support.

How do we measure ROI from AI agents in manufacturing?

ROI should be measured using manufacturing-specific metrics: cycle time reduction, error rate decrease, labor cost savings, throughput improvement, quality improvement, and inventory optimization. The best AI agent development companies help you define success metrics during discovery, baseline current performance, and track improvements throughout implementation. USM provides regular scorecards showing progress toward business outcomes, not just technical milestones.

What if the AI agent makes a mistake in our production environment?

This is why governance is critical and must be built-in from day one. Properly designed AI agents include approval gates for high-stakes decisions, comprehensive audit trails, confidence thresholds (escalating to humans when uncertain), and rollback capabilities. Implementations should start with a pilot phase where the agent runs with human oversight before full autonomous operation. Manufacturing specialists like USM design with these safeguards as core architecture.

 

 

The first 30 days of agentic AI governance: A practical checklist

Every agent you deploy expands your blast radius. A predictive model can produce a bad response, but an agent can act on it.

Agents can retrieve sensitive data, change systems of record, trigger workflows, or pass errors to other agents. The risk is no longer just model quality. It is the authority an agent holds, the systems it can reach, and how quickly a failure can spread.

Eliminating autonomy isn’t the answer. Autonomy without governance creates unmanaged risk. Governance that blocks autonomy creates stagnation. The goal is controlled autonomy: enough authority to create value, with behavior that remains bounded, observable, and interruptible.

CIOs and AI leaders should be able to ask six questions about every production agent and receive clear, evidence-backed answers:

  • Which agent acted?
  • What was it authorized to do?
  • Which data, tools, and systems did it use?
  • Which policies governed the action?
  • Can we reconstruct its actions and reverse-engineer the outcome?
  • Who can intervene right now?

You don’t need to implement every control yourself. But you do need to know what to ask your teams, what “done” looks like, and what risk the organization is accepting when an answer remains unclear.

The first 30 days should establish the controls needed to answer these questions without launching a new investigation. Define the agent. Limit its authority. Track its actions. Test its boundaries. Give someone the power to stop it. Governance will mature over time, but production agents should never operate on trust alone.

Key takeaways

  • Treat every AI agent as a distinct enterprise actor with a named owner, defined purpose, and bounded scope.
  • Give agents only the data, tools, and actions required for that scope. Make access attributable and revocable.
  • Enforce high-impact boundaries through deterministic runtime controls rather than relying on model instructions alone.
  • Record the complete execution path so teams can reconstruct what the agent did and determine why.
  • Test failure conditions as seriously as the happy path, and assign people who can investigate, suspend, and safely restore the agent.
Phase Leadership question What “done” looks like
Days 1–5 Can you identify the agent and its authority? Every agent has a unique identity, owner, bounded scope, and system inventory.
Days 6–10 Can you confirm permissions are enforced at runtime? Every tool and action maps to a defined, attributable, and revocable permission.
Days 11–15 Are high-impact actions governed outside the model? Deterministic controls block, redirect, or escalate actions that violate policy.
Days 16–20 Can your teams reconstruct every consequential action? Teams can trace a complete run from request through downstream effects.
Days 21–25 Does the agent fail safely beyond the happy path? Known failure modes are documented, tested, and reflected in policy thresholds.
Days 26–30 Can named owners stop and restore the agent? Named owners can suspend, investigate, and safely restore the agent.

Days 1–5: Can you identify the agent and its authority?

You can’t govern “the customer service agent” or “the finance copilot” as an informal concept. Every production agent needs a distinct identity and a precise definition of what it’s authorized to do.

Create an agent record that captures:

  • A unique identity, named owner, business purpose, and risk classification
  • The models, tools, APIs, data sources, and downstream systems it uses
  • The actions it may recommend, initiate, approve, or never perform
  • Its escalation boundaries and conditions for human intervention

Specificity is the key. Define the scope in specific, enforceable terms: “Retrieve approved knowledge-base content, summarize account history, and draft responses for human approval.” This gives security, compliance, and engineering teams clear boundaries they can implement and enforce.

Document negative scope, too. Can the agent issue refunds? Change account entitlements? Retrieve payment data? Contact a customer without approval? Unclear answers signal unresolved production risk.

Milestone: Every agent has an identity, owner, explicit action boundary, and inventory of connected resources.

Days 6–10: Can you confirm permissions are enforced at runtime?

An agent’s documented scope matters only if the organization can enforce it when the agent acts.

Identity establishes which actor is operating. Authorization determines what that actor is allowed to do. Apply least-privilege access based on the agent’s assigned task, not the broadest workflow it may eventually support. Separate read, write, execute, and administrative permissions. Permission to retrieve a record should not automatically include permission to modify or delete it.

Apply the strictest authorization requirements to high-impact capabilities, including:

  • Writes to systems of record
  • Financial transactions
  • Access to sensitive data
  • External communications
  • Code execution
  • Tools exposed through Model Context Protocol (MCP) servers or other agent interfaces

Avoid shared service accounts. They obscure attribution and make access reviews unreliable. Use agent-specific credentials, short-lived tokens, conditional access, and explicit tool allowlists where possible.

Define the exception process in advance. Specify who can approve temporary elevation, how long it can remain active, and which actions always require human approval. Authorization should fail closed. If identity or operating context cannot be verified, or an action cannot be evaluated against policy, the agent should stop or escalate rather than improvise.

Milestone: Every tool call is evaluated against defined permissions. Elevated access is conditional and time-bound, and every exception has a designated approver and expiration.

Days 11–15: Are high-impact actions governed outside the model?

This is where controlled autonomy becomes operational: the model can propose an action, but it cannot decide for itself whether that action is permitted.

Permissions and guardrails address different risks. Permissions define what an agent can access. Guardrails constrain how the agent can use that access. Guardrails are enforced through validation, policy checks, and other controls placed throughout the workflow.

Apply policy checks throughout the workflow, not only to the final response. Inspect user inputs, retrieved context, model outputs, tool arguments, and proposed actions for personally identifiable information, prompt injection, unsafe content, policy violations, and prohibited behavior. A final-output review alone does not govern the steps where the agent reads sensitive data, constructs tool calls, or initiates consequential actions.

Prompt instructions such as “never reveal sensitive data” are not sufficient. Malicious or conflicting instructions can enter through user input, retrieved documents, tool output, or another agent. Enforce guardrails at the boundaries between the agent and the resources it can read, modify, or affect.

For high-impact actions, use deterministic policy checks outside the model. Before a tool executes, validate transaction limits, approved recipients, required fields, data classifications, and human approval requirements. The model may propose an action, but the policy layer decides whether the system permits it.

Milestone: Policy checks run before sensitive data crosses a boundary or a high-impact action executes. Failed checks trigger a defined block, fallback, or escalation.

Days 16–20: Can your teams reconstruct every consequential action?

Governance depends on being able to reconstruct what an agent did, why it did it, and what happened next. Final outputs are not enough. Teams need visibility into the full execution path, including the information the agent received, the tools it called, the permissions and policy checks applied, and the actions that affected downstream systems.

Capture the key elements of each run:

  • The original request, system instructions, model version, and policy version
  • Retrieved context, tool calls, permission decisions, and executed actions
  • Downstream effects, human approvals, overrides, and interventions

Use correlation identifiers to connect activity across tools, systems, and agents. Protect logs from tampering, define appropriate retention periods, and limit access to audit data. Logging should improve accountability without creating a new repository of exposed sensitive information.

Operational monitoring should focus on signals that indicate misuse, failure, or drift. Track access violations, abnormal tool activity, repeated retries, latency spikes, cost anomalies, and policy exceptions. Route each signal to a team with the authority and responsibility to investigate. A dashboard without a named owner does not provide meaningful oversight.

Milestone: Security, platform, and compliance teams can reconstruct any consequential agent run from the original request through its downstream effects. Actionable anomaly alerts are routed to named owners.

Days 21–25: Does the agent fail safely beyond the happy path?

The happy path proves that the agent can complete its intended workflow when inputs are clear, data is accurate, tools are available, and policies align. Governance testing must also prove that it fails safely when those conditions break down.

Test ambiguous requests, incomplete records, conflicting policies, unavailable tools, stale data, malicious retrieved content, and attempts to exceed authority. Include multi-step scenarios in which an apparently harmless first action creates risk later in the workflow.

Measure both failure modes: controls that are too weak and controls that are too restrictive. Weak controls create exposure. Overly restrictive controls reduce utility, increase unnecessary escalations, and prevent adoption.

Use early deployments to tune policy thresholds, escalation logic, and intervention triggers. Track task success alongside blocked actions, override rates, false positives, escalation time, and action reversibility.

Milestone: The agent succeeds on representative happy-path workflows, passes adversarial and boundary testing, and has documented failure modes and policy thresholds that reflect an explicit risk-value tradeoff.

Days 26–30: Can named owners stop and restore the agent?

Governance fails when everyone is responsible in principle and no one is accountable in practice.

Name owners for agent performance, access, compliance, monitoring, and incident response. Define who investigates anomalies, who approves remediation, and who has the authority to suspend the agent.

Document rollback, credential revocation, tool isolation, kill switch activation, human takeover, evidence preservation, and post-incident review. Then rehearse the process. A kill switch that has never been tested is only a theory.

Set a review cadence for permissions, policy compliance, operational performance, and business impact. Agent scope, connected tools, and policies will change. Governance must detect that drift before it becomes an incident.

Milestone: Named owners can suspend, investigate, and safely restore the agent through a tested process with clear decision rights.

What operational governance looks like after 30 days

After 30 days, your teams should be able to answer the six questions above with current records and operational evidence. If an answer depends on institutional memory or an unmaintained spreadsheet, the control is not operational.

This isn’t a complete governance program. It is the foundation for one. Start by making each agent legible, bounded, observable, and interruptible. As your agent footprint expands, these controls will require centralized automation.

Autonomy without governance creates unmanaged risk. Governance that blocks autonomy creates stagnation. The first 30 days establish the middle path: controlled autonomy that can earn trust and scale.

For the complete framework, download The enterprise guide to agentic AI governance.

The post The first 30 days of agentic AI governance: A practical checklist appeared first on DataRobot.

The first 30 days of agentic AI governance: A practical checklist

Every agent you deploy expands your blast radius. A predictive model can produce a bad response, but an agent can act on it.

Agents can retrieve sensitive data, change systems of record, trigger workflows, or pass errors to other agents. The risk is no longer just model quality. It is the authority an agent holds, the systems it can reach, and how quickly a failure can spread.

Eliminating autonomy isn’t the answer. Autonomy without governance creates unmanaged risk. Governance that blocks autonomy creates stagnation. The goal is controlled autonomy: enough authority to create value, with behavior that remains bounded, observable, and interruptible.

CIOs and AI leaders should be able to ask six questions about every production agent and receive clear, evidence-backed answers:

  • Which agent acted?
  • What was it authorized to do?
  • Which data, tools, and systems did it use?
  • Which policies governed the action?
  • Can we reconstruct its actions and reverse-engineer the outcome?
  • Who can intervene right now?

You don’t need to implement every control yourself. But you do need to know what to ask your teams, what “done” looks like, and what risk the organization is accepting when an answer remains unclear.

The first 30 days should establish the controls needed to answer these questions without launching a new investigation. Define the agent. Limit its authority. Track its actions. Test its boundaries. Give someone the power to stop it. Governance will mature over time, but production agents should never operate on trust alone.

Key takeaways

  • Treat every AI agent as a distinct enterprise actor with a named owner, defined purpose, and bounded scope.
  • Give agents only the data, tools, and actions required for that scope. Make access attributable and revocable.
  • Enforce high-impact boundaries through deterministic runtime controls rather than relying on model instructions alone.
  • Record the complete execution path so teams can reconstruct what the agent did and determine why.
  • Test failure conditions as seriously as the happy path, and assign people who can investigate, suspend, and safely restore the agent.
Phase Leadership question What “done” looks like
Days 1–5 Can you identify the agent and its authority? Every agent has a unique identity, owner, bounded scope, and system inventory.
Days 6–10 Can you confirm permissions are enforced at runtime? Every tool and action maps to a defined, attributable, and revocable permission.
Days 11–15 Are high-impact actions governed outside the model? Deterministic controls block, redirect, or escalate actions that violate policy.
Days 16–20 Can your teams reconstruct every consequential action? Teams can trace a complete run from request through downstream effects.
Days 21–25 Does the agent fail safely beyond the happy path? Known failure modes are documented, tested, and reflected in policy thresholds.
Days 26–30 Can named owners stop and restore the agent? Named owners can suspend, investigate, and safely restore the agent.

Days 1–5: Can you identify the agent and its authority?

You can’t govern “the customer service agent” or “the finance copilot” as an informal concept. Every production agent needs a distinct identity and a precise definition of what it’s authorized to do.

Create an agent record that captures:

  • A unique identity, named owner, business purpose, and risk classification
  • The models, tools, APIs, data sources, and downstream systems it uses
  • The actions it may recommend, initiate, approve, or never perform
  • Its escalation boundaries and conditions for human intervention

Specificity is the key. Define the scope in specific, enforceable terms: “Retrieve approved knowledge-base content, summarize account history, and draft responses for human approval.” This gives security, compliance, and engineering teams clear boundaries they can implement and enforce.

Document negative scope, too. Can the agent issue refunds? Change account entitlements? Retrieve payment data? Contact a customer without approval? Unclear answers signal unresolved production risk.

Milestone: Every agent has an identity, owner, explicit action boundary, and inventory of connected resources.

Days 6–10: Can you confirm permissions are enforced at runtime?

An agent’s documented scope matters only if the organization can enforce it when the agent acts.

Identity establishes which actor is operating. Authorization determines what that actor is allowed to do. Apply least-privilege access based on the agent’s assigned task, not the broadest workflow it may eventually support. Separate read, write, execute, and administrative permissions. Permission to retrieve a record should not automatically include permission to modify or delete it.

Apply the strictest authorization requirements to high-impact capabilities, including:

  • Writes to systems of record
  • Financial transactions
  • Access to sensitive data
  • External communications
  • Code execution
  • Tools exposed through Model Context Protocol (MCP) servers or other agent interfaces

Avoid shared service accounts. They obscure attribution and make access reviews unreliable. Use agent-specific credentials, short-lived tokens, conditional access, and explicit tool allowlists where possible.

Define the exception process in advance. Specify who can approve temporary elevation, how long it can remain active, and which actions always require human approval. Authorization should fail closed. If identity or operating context cannot be verified, or an action cannot be evaluated against policy, the agent should stop or escalate rather than improvise.

Milestone: Every tool call is evaluated against defined permissions. Elevated access is conditional and time-bound, and every exception has a designated approver and expiration.

Days 11–15: Are high-impact actions governed outside the model?

This is where controlled autonomy becomes operational: the model can propose an action, but it cannot decide for itself whether that action is permitted.

Permissions and guardrails address different risks. Permissions define what an agent can access. Guardrails constrain how the agent can use that access. Guardrails are enforced through validation, policy checks, and other controls placed throughout the workflow.

Apply policy checks throughout the workflow, not only to the final response. Inspect user inputs, retrieved context, model outputs, tool arguments, and proposed actions for personally identifiable information, prompt injection, unsafe content, policy violations, and prohibited behavior. A final-output review alone does not govern the steps where the agent reads sensitive data, constructs tool calls, or initiates consequential actions.

Prompt instructions such as “never reveal sensitive data” are not sufficient. Malicious or conflicting instructions can enter through user input, retrieved documents, tool output, or another agent. Enforce guardrails at the boundaries between the agent and the resources it can read, modify, or affect.

For high-impact actions, use deterministic policy checks outside the model. Before a tool executes, validate transaction limits, approved recipients, required fields, data classifications, and human approval requirements. The model may propose an action, but the policy layer decides whether the system permits it.

Milestone: Policy checks run before sensitive data crosses a boundary or a high-impact action executes. Failed checks trigger a defined block, fallback, or escalation.

Days 16–20: Can your teams reconstruct every consequential action?

Governance depends on being able to reconstruct what an agent did, why it did it, and what happened next. Final outputs are not enough. Teams need visibility into the full execution path, including the information the agent received, the tools it called, the permissions and policy checks applied, and the actions that affected downstream systems.

Capture the key elements of each run:

  • The original request, system instructions, model version, and policy version
  • Retrieved context, tool calls, permission decisions, and executed actions
  • Downstream effects, human approvals, overrides, and interventions

Use correlation identifiers to connect activity across tools, systems, and agents. Protect logs from tampering, define appropriate retention periods, and limit access to audit data. Logging should improve accountability without creating a new repository of exposed sensitive information.

Operational monitoring should focus on signals that indicate misuse, failure, or drift. Track access violations, abnormal tool activity, repeated retries, latency spikes, cost anomalies, and policy exceptions. Route each signal to a team with the authority and responsibility to investigate. A dashboard without a named owner does not provide meaningful oversight.

Milestone: Security, platform, and compliance teams can reconstruct any consequential agent run from the original request through its downstream effects. Actionable anomaly alerts are routed to named owners.

Days 21–25: Does the agent fail safely beyond the happy path?

The happy path proves that the agent can complete its intended workflow when inputs are clear, data is accurate, tools are available, and policies align. Governance testing must also prove that it fails safely when those conditions break down.

Test ambiguous requests, incomplete records, conflicting policies, unavailable tools, stale data, malicious retrieved content, and attempts to exceed authority. Include multi-step scenarios in which an apparently harmless first action creates risk later in the workflow.

Measure both failure modes: controls that are too weak and controls that are too restrictive. Weak controls create exposure. Overly restrictive controls reduce utility, increase unnecessary escalations, and prevent adoption.

Use early deployments to tune policy thresholds, escalation logic, and intervention triggers. Track task success alongside blocked actions, override rates, false positives, escalation time, and action reversibility.

Milestone: The agent succeeds on representative happy-path workflows, passes adversarial and boundary testing, and has documented failure modes and policy thresholds that reflect an explicit risk-value tradeoff.

Days 26–30: Can named owners stop and restore the agent?

Governance fails when everyone is responsible in principle and no one is accountable in practice.

Name owners for agent performance, access, compliance, monitoring, and incident response. Define who investigates anomalies, who approves remediation, and who has the authority to suspend the agent.

Document rollback, credential revocation, tool isolation, kill switch activation, human takeover, evidence preservation, and post-incident review. Then rehearse the process. A kill switch that has never been tested is only a theory.

Set a review cadence for permissions, policy compliance, operational performance, and business impact. Agent scope, connected tools, and policies will change. Governance must detect that drift before it becomes an incident.

Milestone: Named owners can suspend, investigate, and safely restore the agent through a tested process with clear decision rights.

What operational governance looks like after 30 days

After 30 days, your teams should be able to answer the six questions above with current records and operational evidence. If an answer depends on institutional memory or an unmaintained spreadsheet, the control is not operational.

This isn’t a complete governance program. It is the foundation for one. Start by making each agent legible, bounded, observable, and interruptible. As your agent footprint expands, these controls will require centralized automation.

Autonomy without governance creates unmanaged risk. Governance that blocks autonomy creates stagnation. The first 30 days establish the middle path: controlled autonomy that can earn trust and scale.

For the complete framework, download The enterprise guide to agentic AI governance.

The post The first 30 days of agentic AI governance: A practical checklist appeared first on DataRobot.

Surviving the paper deluge: a one-year study in learning from demonstration

With the explosion of robotics research, staying current in fields like Learning from Demonstration (LfD) is a monumental challenge. Is AI the solution to the “paper deluge,” or is it part of the problem? Read the article preview below to learn more!

Download the full paper: Surviving the Paper Deluge.

Authors: Aude Billard, Renaud Detry, Nadia Figueroa, Maximilian Foriest, Dongheui Lee, Kunpeng Yao
Contributions: The five senior authors (A.B, R.D, N.F, D.L and K.Yao) collectively designed the study, read the papers, conducted the qualitative and quantitative analysis and writing of the paper. M. F. contributed scripts for LLM analysis and participated in LLM-Human comparison.


Summary

Scientists are expected to read newly published papers in their field to stay current and keep their work relevant. However, when faced with the massive number of publications, it may seem an overwhelming task to read all these papers, even if one were to reduce this to only a fraction related to one’s own area of research. As an example, in 2024 alone, IEEE published no less than 46,968 papers on “robotics” or “automation”, and IEEE publications represent only a fraction of the total research available online

To assess the magnitude of this challenge, as well as to evaluate how much genuine progress is reported in today’s publications, we undertook exactly this effort. For the task to be reasonable, we reduced our search to one particular subarea, learning from demonstration (LfD), that is methods whereby robots are taught by human experts. We monitor progress through both quantitative and qualitative metrics, offering a review on current trends and notable contributions. We also delineate areas of importance, but that seem to receive little attention and offer recommendations for promoting.

Our assessment was primarily based both on a human-eye assessment of all papers. We also explored the use of AI and other computing tools to do this task in our place. While scripts and large language models (LLMs) can be used fairly faithfully to provide general quantitative assessment, they fail when it comes to assessing the true importance of the research. They cannot recognize a paper revisiting a work that already had solutions. They fail to recognize when the abstract or claims of the paper are overstatements over the true contribution reported in the paper.

Our overall assessment led us to conclude that from a deck of more than 300 papers, only about 20% of the papers could be qualified as offering highly notable contributions, while the remainder of the papers offered a variety of incremental improvements over existing methods, or new domains of applications. The notable contributions did not correlate necessarily with a higher number of downloads or citations. Finding these gems is, however, essential to reduce the risk that novel work goes unnoticed and reduce duplication of efforts. We offer a few thoughts on how to best combine direct reading of the literature with automated approaches (scripts and LLMs) to streamline the review process. We close with a few recommendations: a) develop a research engine that restores the natural importance of work done by journal and conference editorial boards to rank papers based on evaluation scores and peer-reviewed status, in place of Google Scholar or IEEEXplore, that place all publications on equal footing, disregarding peer reviewing and the reputation of journals and conferences, b) consider establishing a blind publication model and topic-based social media posting, where authors’ name and institution are downplayed and become accessory to the paper to ensure that focus be on the content of the publication rather than secondary aspects, c) take a holistic approach to use of LLM in support of reviewing literature, using them for what they excel at, namely summarizing a piece of work and collecting precise quantitative information, but bearing in mind that, while today the tools cannot match expert capacity to assess true novelty, should they achieve this one day, this may have repercussion on our own ability to provide said expertise.

Publications growth

Over the past decade, the number of submissions to robotics journals has grown steadily on a yearly basis, with an explosive trend in 2023 (26%) and 2024 (31%), likely due to different factors, including growing interest in the public and private sectors and to the availability of AI tools supporting the writing of papers and code. The number of published papers has closely followed this trend, despite all efforts made by editorial boards to contain the growth by decreasing acceptance rates. Conferences have followed the same trend. For instance, ICRA doubled the number of papers it published in ten years, reaching approximately 1,800 in 2024. Simultaneously, the strong pressure exerted by the community to publish rapidly has led to a 50% decrease in the time window between the submission of a paper and its publication. The phenomenon is not particular to IEEE publications, and journals and conferences such as IJRR, RSS and CoRL have followed the same trend.

Clearly, it would be unrealistic to expect any researcher to read all of these publications. One might argue that researchers are typically interested in only a subset of the literature, for instance a specific domain or methodology, and would therefore read only a fraction of all published papers. Yet even this narrower scope may prove unmanageable. To assess how feasible it is for a researcher to stay current within their own area of expertise, we undertook the task of reading a large fraction of all papers published in our domain – learning from demonstration – over the course of a single year (2024).


This article originally appeared on IEEE RAS.

Page 1 of 8
1 2 3 8