Home

  • Meta launches Muse Code: A new AI coding agent to rival OpenAI and Anthropic

    • Muse Code can handle complete software engineering tasks across large repositories, making coding more accessible.
    • Meta aims to differentiate Muse Code by offering more affordable pricing tiers than competitors like Claude and Codex.
    • The tool is powered by an updated Muse Spark model, which allows for parallel processing of coding tasks without collisions.

    [via]

  • Synthetic consumers can produce answers. But can you bet a product launch on them?

    Synthetic consumers can produce answers. But can you bet a product launch on them?

    Synthetic consumer research is becoming one of the most funded ideas in AI. Companies like Simile are creating digital twins of people that aim to predict how consumers will respond to products, prices, features, and marketing messages.

    The appeal is clear: Instead of taking weeks to recruit participants, conduct interviews, and analyze responses, a company can generate hundreds or thousands of synthetic reactions in a matter of hours.

    This makes synthetic research faster, cheaper, and significantly easier to scale without needing to talk to a single customer.

    But speed does not equal evidence.

    The main issue with synthetic consumers is not that their answers are always wrong. It is that businesses often cannot tell when those answers are incorrect. This distinction is crucial when the research is used to make costly, irreversible decisions.

    Synthetic answers lack real human evidence.

    Most synthetic research systems combine large language models with demographic information, behavioral datasets, social media content, transaction data, or prior research. The model then produces responses that mimic what someone from a specific segment might say.

    However, the answer still originates from a model.

    There is no real customer behind the statement. There is no respondent whose situation can be examined. No interview recording exists to revisit, no behavior to observe, and no person for the researcher to question further.

    A synthetic respondent might say:

    “I would pay ₹2,499 for a premium protein supplement because I value clean ingredients.”

    That sounds helpful. But what does it prove?

    It does not show whether a real buyer will pay ₹2,499 when a competing product is available for ₹1,799. It does not reveal if the buyer will abandon the cart after seeing the delivery fee. It does not take into account advice from a gym trainer, distrust of an unfamiliar brand, concerns about taste, or a spouse questioning the monthly expense.

    The model has created a plausible explanation, but it has not observed a purchase.

    This makes synthetic data hard to use as evidence for high-stakes decisions. A company cannot confidently tell its board, product team, or investors that customers demanded a feature when no actual customers contributed to that finding.

    Plausibility is not the same as prediction.

    Large language models are good at generating answers that a reasonable person might give.

    Unfortunately, humans are not always reasonable.

    People contradict themselves. They forget why they bought something. They claim to care about sustainability but then choose the product that arrives tomorrow. They say they want fewer notifications but continue opening apps designed for notifications. They demand privacy yet trade personal information for a small discount.

    Research has shown that synthetic respondents may show less variation than real consumers and can exaggerate the relationship between demographics and attitudes. In other words, they can make consumer segments seem more consistent internally and more different from each other than real people actually are.

    This is a structural problem, not just an accuracy problem.

    A synthetic consumer is built from patterns. A real consumer is shaped by constraints, relationships, habits, contradictions, and moments of irrationality. These elements often drive the purchase.

    The melody incident serves as a cautionary tale.

    In May 2026, a video featuring Indian Prime Minister Narendra Modi, Italian Prime Minister Giorgia Meloni, and Melody toffees gained attention on social media.

    Retail investors then bought shares of Parle Industries, mistakenly associating the listed company with Melody. However, Melody is made by Parle Products, a separate and privately held company. Parle Industries had no ties to the chocolate.

    The sequence was irrational but recognizable as human behavior:

    • PM Modi gifts Melody to Italian PM.
    • Parle makes Melody.
    • A listed company has “Parle” in its name.
    • Buy the stock.

    Similar errors have occurred elsewhere. After Elon Musk tweeted “Use Signal,” investors inflated shares of Signal Advance, an unrelated medical device company. Investors also repeatedly confused Zoom Video Communications with the unrelated Zoom Technologies.

    These are extreme cases, but they highlight the core problem.

    Human decisions are influenced by availability bias, mistaken connections, social proof, the fear of missing out, and whatever stands out at that moment.

    A model trained to create coherent behavior may consistently overlook incoherent behavior.

    Consider a protein brand deciding on its next flavor.

    Suppose a protein company must decide whether to launch mango, chocolate hazelnut, or unflavored whey.

    A synthetic panel can analyze category trends, reviews, demographic preferences, and social media discussions. It might conclude that mango will attract young Indian consumers because it is familiar, culturally relevant, and different from existing chocolate products.

    That is a reasonable guess.

    But the actual purchase may rely on factors the model cannot grasp:

    • Does mango whey taste artificial when mixed with water?
    • Does its smell become unpleasant after resting in a shaker for an hour?
    • Do consumers think of mango as a refreshing drink instead of a heavy protein product?
    • Will gym trainers recommend it?
    • Does the bright packaging make it appear less serious than competing products?
    • Will customers enjoy the first serving but tire of the flavor after ten days?

    These are not just data points. They are experiences.

    A synthetic consumer can describe what consuming mango protein might feel like. It cannot actually taste it repeatedly, grow bored with it, regret buying a one-kilogram pack, or leave the half-used container at the back of a kitchen shelf.

    This distinction matters because the company is not deciding which idea sounds best. It is deciding which product to manufacture, stock, distribute, and promote.

    Where else can synthetic research lead to mistakes?

    Consider pricing.

    A synthetic respondent may weigh price, ingredients, and brand reputation rationally. A real buyer may choose the most expensive product because a fitness influencer recommends it or the cheapest one because payday is still a week away.

    Consider packaging.

    An AI persona can evaluate the visual design presented on a screen. It cannot discover that the lid is hard to open, the scoop gets buried in the powder, or the container doesn’t fit in a kitchen cabinet.

    Consider customer churn.

    A model may suggest that customers cancel due to high prices. Interviews might reveal that customers actually felt embarrassed asking the support team the same question repeatedly or that a spouse objected to another subscription showing up on the credit card statement.

    Consider a new feature.

    Synthetic users may consistently prefer more control and customization. Real users may never figure out the settings, may feel overwhelmed, or may stick with the default because changing their habits takes effort.

    Consider advertising.

    A synthetic audience may understand the intended message. Real consumers may misinterpret one line, turn a screenshot into a meme, or link the campaign with a controversy that did not exist when the research data was gathered.

    Synthetic research is weakest where businesses most need it: where context, behavior, and consequences matter.

    The accuracy debate is a distraction.

    Synthetic research companies often discuss whether their predictions are 80%, 90%, or 95% accurate.

    But an average accuracy number gives decision-makers little insight.

    Accurate at predicting what?

    Under what conditions?

    For which group of people?

    Compared to which human sample?

    Is the data reliable?

    Does the system predict average survey responses, individual choices, market share, repeat purchases, or real behavior under financial pressure?

    A system might mirror broad consumer sentiment with 90% accuracy yet still fail at predicting the minority behaviors crucial for a particular product’s success. It could correctly identify chocolate as the safest flavor but miss the small but valuable group willing to pay significantly more for an unflavored, clean-label item.

    Even a highly accurate model can pose risks when users do not understand where the remaining errors lie.

    The issue is not whether synthetic data can resemble human responses. It clearly can. The issue is whether that resemblance stays reliable when the market shifts, the product is new, or the decision relies on an unusual human reaction.

    These are often the very situations where companies seek research.

    Synthetic data is useful—but only to a point.

    Synthetic research should not be completely dismissed.

    It can assist teams in forming hypotheses, exploring potential segments, stress-testing questionnaires, identifying obvious objections, and narrowing down a large array of concepts before engaging with customers. It can also guide researchers in deciding which questions need deeper exploration.

    In these cases, synthetic data serves as a brainstorming tool rather than as proof.

    The problem arises when generated responses are presented as customer evidence.

    There is a big difference between saying:

    “The simulation suggests that price may be a concern.”

    and saying:

    “Our customers told us that price is the primary barrier.”

    Only the second statement requires actual customers.

    A reasonable research process can use synthetic consumers early on to explore possibilities. However, before making decisions about product launches, pricing, positioning, inventory, or major investments, those possibilities must be tested with real people and, whenever feasible, actual behavior.

    The question is not whether synthetic consumers are impressive. The question is whether a company should produce ten thousand units, change its pricing, or enter a new market.

    For low-risk exploration, synthetic data may be sufficient.

    For decisions where being wrong is expensive, plausible answers are not enough. Businesses need evidence that can be traced to real people, real circumstances and real decisions.

    Synthetic consumers can tell you what might happen.

    (Human) Customer research tells you what people are actually experiencing.

    And behavioural data tells you what they ultimately did.


    What’s your take?

  • AI automation leads to reduced entry-level hiring at 22% of companies

    • A recent survey by Gartner reveals that 22% of Chief Human Resource Officers (CHROs) have observed a halt in entry-level hiring due to the rise of AI automation.
    • This trend indicates a significant shift in workforce dynamics, as companies leverage technology to streamline operations.
    • The impact could lead to fewer job opportunities for new entrants in the labor market, raising concerns about workforce development.

    [via]

  • LinkedIn lets users flag AI-generated content

    • LinkedIn is testing a feature allowing users to report posts and comments as ‘AI slop’.
    • The initiative aims to improve content quality and combat the rise of automated comments.
    • Users will receive feedback on reported posts to help refine their use of AI tools.

    [via]

  • Alibaba launches powerful Qwen3.8-Max LLM with 2.4 trillion parameters

    • Qwen3.8-Max features 2.4 trillion parameters, making it Alibaba’s most advanced LLM to date.
    • The model can process prompts of up to 1 million tokens, analyzing extensive text and video data.
    • Qwen3.8-Max demonstrated its capabilities by completing complex coding and chip design tasks autonomously.

    [via]

  • GenOffice

    GenOffice

    GenOffice — An AI-native office suite for macOS and Windows

    GenOffice is an AI-native office suite that includes a word processor, spreadsheet, presentations, and PDF tools.

    • Available for both macOS and Windows platforms.
    • Includes essential office applications like word processor, spreadsheet, and presentation tools.
    • Utilizes AI technology to enhance productivity and user experience.

    [Get it]

  • pdf-inspector (open source)

    pdf-inspector — Fast Rust library for PDF inspection, classification, and text extraction.

    pdf-inspector is a Rust library designed to inspect, classify, and extract text from PDF documents, intelligently distinguishing between scanned and text-based PDFs.

    • Intelligently detects scanned vs text-based PDFs for smart routing decisions.
    • Supports text extraction and classification of PDF documents.
    • Provides browser bindings for WebAssembly integration.

    [Get it]

  • agentic crm (open source)

    agentic crm (open source)

    crm — An agentic-first system for managing customer relationships.

    crm is a customer relationship management tool designed to facilitate AI-native applications and enhance data handling.

    • Features a library of AI Elements components for building intelligent applications.
    • Includes robust task management and error handling for agent scheduling.
    • Provides clear documentation and guidelines for developers and users.

    [Get it]

  • qm

    qm

    qm — Multiplayer agent harness for work.

    qm is a framework designed to facilitate the development and deployment of multiplayer agent-based applications.

    • Supports integration with various AI models and providers.
    • Provides a CLI for easy setup and management of agent deployments.
    • Includes features for handling authentication and external OIDC providers.

    [Get it]

  • Sarvam AI: Here is all that was announced, claimed and debated

    Sarvam AI: Here is all that was announced, claimed and debated

    Bengaluru got another big AI day this week. Sarvam used Epoch to push its full-stack story harder: trillion-parameter model in the works, upgrades to the 105B, better speech models, Vision 2.0, coding agents, India-hosted inference, even more smartglasses demos. Ambitious, loud, and very on-brand for a company that has positioned itself as India’s main sovereign AI bet.

    The reaction has been a mix of real interest and quiet eye-rolling.

    The claims in short: they’re building a trillion-plus parameter model from scratch in India, focused on coding, cybersecurity, science and simulation. Roughly six-month timeline according to the messaging.

    Current 105B is being sold as roughly $0.80 per million blended tokens, which they say is 5.5 times cheaper than GPT-5.4 Mini and about 11 times cheaper than Gemini 3.5 Flash. Voice is the big flex. They keep saying if you’re building voice products for India right now, nothing is cheaper or more scalable (we do have questions on latency).

    Speech is where people actually perked up.

    Saras V4 (speech-to-text) claims better coverage of lower-resource Indian languages and competitive English numbers. Bulbul V4 (text-to-speech) adds emotion and naturalness; several builders said the Hindi output is among the best they’ve heard. Vision 2.0 improves OCR on Indian handwriting and documents. There’s also Sarvam Code, local inference options, telephony tools, and talk of scaling Blackwell clusters plus a San Francisco office.

    Pricing for people who actually ship:

    • 105B: ₹4 input / ₹2.5 cached / ₹16 output per million tokens
    • 30B is cheaper
    • Speech-to-text: ₹30 per hour (₹45 with diarization)
    • Text-to-speech: ₹15-30 per 10k characters depending on version
    • Vision: ₹0.5 per page

    If the quality holds in production, the economics for call centres, BFSI and government work look interesting. That’s the real hook.

    What people are actually questioning:

    How much of the coding and agent stuff is real model progress versus a polished harness around other models? GLM keeps coming up in the side conversations. Were the benchmark slides selective? A few people noticed missing or conveniently ranked competitors.

    Is Sarvam still trying to be a frontier model lab, or has it become an applied AI and infrastructure company that also trains models? The trillion-parameter plan sounds good on stage. Can they actually train and serve something competitive on the timelines and hardware they have, or does this become another ambitious slide?

    Why aren’t more Indian product companies already deep on the 105B if the cost story is this strong? Latency, reliability at scale and basic ecosystem maturity still come up in private chats. After the capital raised and the government proximity, is the delivery matching the narrative?

    One post that landed with people: the speech work is solid and didn’t need the questionable comparison graphs. Don’t spend the goodwill on theatre.

    Sarvam occupies an important spot.

    After other Indian efforts shifted focus, a lot of builders still want this one to work. The full-stack bet (models + speech + vision + inference + agents, India-first) is coherent.

    Cost advantages in voice and local languages are not trivial. Sovereignty messaging hits differently when the alternative is shipping every conversation overseas.

    At the same time the Indian AI conversation has grown up. People now ask harder questions about evaluation honesty, actual capability versus packaging, and whether capital is turning into durable technical edge. That’s healthy.

    Epoch was neither a disaster nor a coronation. It was a serious company showing its current hand while reaching for a much bigger one. The next independent evaluations of the 105B, real production use of the voice stack, and visible progress on the trillion-parameter effort will matter more than any conference day.

    Until then the questions stay open. As they should.