Agents for financial services
Anthropic releases ten Cowork and Claude Code plugins plus Microsoft 365 integrations and MCP app for financial services.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Anthropic releases ten Cowork and Claude Code plugins plus Microsoft 365 integrations and MCP app for financial services.
The automotive cockpit is undergoing a fundamental shift from rule-based interfaces to agentic, multimodal AI systems capable of reasoning, planning, and... The automotive cockpit is undergoing a fundamental shift from rule-based interfaces to agentic, multimodal AI systems capable of reasoning, planning, and acting. In most vehicles on the road today, in-vehicle assistants still rely on fixed command-response patterns: interpret a phrase, trigger an action, reset. While effective for well-defined tasks, this approach doesn’t scale to modern… Source
Argues frontier AI failures in open-ended tasks (scientific assistance, agents, personalization) stem from objective ambiguity rather than capability gaps; proposes contextual multi-objective optimization.
Generative AI’s explosive first chapter was defined by humans sending requests and models responding. The agentic chapter is different. Agents don't... Generative AI’s explosive first chapter was defined by humans sending requests and models responding. The agentic chapter is different. Agents don’t follow a pre-determined sequence of actions. They call tools, spawn sub-agents with different tasks and models, retain information in memory, manage their own context window, and decide for themselves when they’re finished. In doing so… Source
ProgramBench: 200-task evaluation showing agents struggle to rebuild large binaries from scratch without cheating vulnerabilities.
The Seattle-based startup's Series A round was led by Glilot Capital, NFX and SignalFire, TechCrunch has exclusively learned.
OpenAI and PwC partner to deploy AI agents for enterprise finance automation, forecasting, and CFO workflow modernization.
FlexSQL agent flexibly explores schemas and data during text-to-SQL generation, enabling recovery from early mistakes.
Opinion piece argues LLM agents require jointly formulated actions and plans with human actors rather than isolated architectures.
ORPilot: open-source agentic system translating ambiguous business problems into solver-ready optimization models with conversational and data collection agents.
Zero-trust authorization framework for LLM agents with hybrid inspection and task-based access control to mitigate tool-use and resource-access risks.
Novel framework for bandwidth-efficient remote control via minimal information transmission between controller and agents in continuous action spaces.
DataEvolver implements closed-loop agent-driven visual data generation and refinement for image editing, supporting masks, depth, poses, and trajectory artifacts.
Momentum integrates runtime procedural content generation and autonomous agent evaluation in endless-runner gameplay to assess generated terrain balance and solvability.
Iterated negotiation benchmark tests LLM agents' ability to repair grounding failures in dynamic multi-turn interaction.
GRAVITY module injects relational, temporal, and thematic structure into conversational memory retrieval for long-horizon agents.
AutoMat benchmark evaluates LLM agents on reproducing computational materials science findings, requiring domain knowledge and result interpretation beyond code quality.
MemCoE: cognition-inspired two-stage memory optimization for LLM agents to learn personalized long-term user preferences within context windows.
Analysis of agentic AI specialization: coding agents (Codex-style) for knowledge work, Claude for creative tasks; discusses agents escaping operational boundaries.
Intern-Atlas introduces structured methodological evolution graphs as research infrastructure for AI agents to navigate scientific knowledge beyond citation links.
Link lets users connect cards, banks, and subscriptions, then authorize AI agents to spend securely via approval flows.
STEF enables schema-agnostic evaluation of text-to-SQL agents in production without ground-truth queries, addressing real-world deployment gaps.
Study finds persona prompting in multimodal LLMs produces stable but limited behavioral variation in urban sentiment judgment tasks.
CARE methodology systematizes LLM agent engineering in scientific domains via three-party collaboration between SMEs, developers, and helper agents.
NVIDIA CUDA Tile (cuTile) is a tile-based programming model that enables developers to write GPU kernels in terms of tile-level operations—loads, stores, and... NVIDIA CUDA Tile (cuTile) is a tile-based programming model that enables developers to write GPU kernels in terms of tile-level operations—loads, stores, and matrix multiply-accumulate—rather than manually coordinating threads, warps, and shared memory. cuTile.jl brings the same tile-based approach to the dynamic programming language Julia. Users can write custom GPU kernels without dropping… Source
Architectural pattern language for vision language agents balances latency/non-determinism of VLMs against real-time enterprise control requirements.
Comparative evaluation of three LLM agent paradigms (domain-specific, computer-use, coding) on scientific visualization tasks across 15 benchmarks.
D3-Gym dataset: 565 verifiable tasks from real scientific repositories for evaluating LLM agents on data-driven discovery.
Language model agents design mechanical linkages via symbolic lifting: LLMs explore topologies while numerical optimizers tune parameters, validated on six motion targets.
Survey of RL+GUI agents for long-horizon automation in visual interfaces; proposes framework toward autonomous digital inhabitants with safe exploration.