Run Two AI Models at Once, and Keep the Best Answer, Privately

MTAIChat is the MT Labs self-hosted AI chat platform. Its multi-model mode runs two LLMs side by side on the same prompt so your team can compare answers and keep the better one. It bundles a canvas notes sidebar and unlimited in-chat image generation, runs on a private AI server you rent, and supports your whole company on a flat cost.

The problem: one model is rarely the right answer for everything

Most teams that adopt AI chat settle on a single cloud model and stick with it. That model is excellent at some tasks, mediocre at others, and quietly wrong about a handful of things every week. The team learns its quirks, works around them, and a slow tax builds up across the year.

The honest move is to use more than one model. A reasoning-heavy task wants a model tuned for careful step-by-step work. A drafting task wants a model with a lighter, friendlier register. A summary of a long document wants something with strong recall over a wide context window. No single model is best at all three.

Doing this on per-seat cloud AI chat subscriptions gets expensive fast. Adding a second provider is another seat per user per month, billed forever, with the company name on two invoices instead of one. For a team of fifteen people that is fifteen new line items added to the budget the moment anyone wants to compare.

The quieter problem is where the conversations actually go. Every prompt your team types into a cloud AI service is a conversation with someone else’s infrastructure. For a Singapore SME that handles client data, internal financials, or staff matters, that is a question the PDPA officer will eventually ask, and the answer is usually awkward.

The approach: a multi-model chat on a server you own

MTAIChat is the MT Labs self-hosted AI chat platform. It is built on Open WebUI and runs on the private AI server your team rents from us. The same server that hosts your CRM and your image generator also hosts the chat platform, so the data and the model weights both sit on infrastructure you control.

The headline feature is multi-model mode. You type a prompt and the platform sends it to two models at the same time. The responses appear side by side. Your team reads both, picks the better answer, blends the two, or uses one as a draft and the other as a critique. The compute happens locally, so running a second model is hardware time, not a second subscription.

A few specifics that matter for day to day work:

  • Canvas notes on a right sidebar. Capture working notes, prompt fragments, and partial outputs alongside the chat without losing your place. Useful when a thread is solving a real problem rather than answering a quick question.
  • Unlimited in-chat image generation. The platform is wired to the same local ComfyUI image backbone that runs SecondBrain. Type a prompt for an image and it generates one directly in the chat, with no per-image cloud cost and no separate tool to open.
  • Hands-free voice and video chat. For times when typing is the slow part of the conversation.
  • Granular roles and permissions. An admin panel controls who can do what. Useful when finance, ops, and the creative team all share the platform but need different defaults.
  • Markdown, LaTeX, and a PWA mobile app. Familiar formatting for technical work, and a phone-installable client for staff who are not at a desk.
  • Prefilled prompt suggestions. Non-technical staff get a running start instead of a blank box.

Underneath, the models are local or OpenAI-compatible. Practical choices include Mistral, Qwen, and Gemma in their current generations, with a model builder for custom agents and characters tuned to your own work. Because it is your server, you choose which two models load for the comparison, and you can swap them as better releases ship.

The pricing pattern is the same as the rest of the MT Labs bundle. Unlimited users on a flat server cost. Adding a new hire to MTAIChat does not change the bill, because the chat platform is bundled on the private AI server, not licensed per seat. Per-seat AI chat subscriptions charge for every user every month, and a second model usually means a second subscription on top.

For Singapore companies, the PDPA story sits quietly in the background. Prompts, files, and chat history all stay on the server you rent, inside Singapore, on infrastructure you control. There is no external processor reading the messages, which removes one whole category of cross-border data transfer concern.

What it looks like in practice

A Singapore SME with eighteen people moves its AI chat usage off a pair of per-seat cloud subscriptions. The combined cost was scaling with every new hire, and the data path was making the head of operations uncomfortable.

The new setup is MTAIChat on a rented private server. Two models are loaded for multi-model mode. One is tuned for careful reasoning and the other for fluent writing. The team types a question once and reads two answers, which has become the normal way they use the platform.

The marketing team uses the canvas sidebar for draft fragments while the chat handles research and rewriting. The creative team generates campaign visuals inside the chat using prompts that pull from the same image backbone the design team already uses on SecondBrain, so the look stays consistent across tools. The finance team uses a stricter custom agent with permissions that only allow internal documents as RAG sources.

The platform is also where the MindTheory holiday camps run their student sessions. Kids and parents use MTAIChat hands on at the camps, which is the practical proof that the interface is approachable for first-time users. It is the same product, sized to a different team.

A few months in, the company hires four more people across two departments. The MTAIChat bill does not change. The same hires would have added two cloud chat seats each, every month, on two different subscriptions, indefinitely. The model comparisons keep happening, and the better answer keeps getting used.

Where this approach has limits

A multi-model self-hosted chat is not the right tool for every situation.

If your team is two or three people running a light workload with no strong privacy requirement, a single cloud chat subscription will still feel like the lower friction option. The bundle pays off when there is a team to consolidate and a real reason to keep the conversations on local infrastructure.

If your work depends on a very specific frontier model that is only available through one vendor, the honest answer is that an open or compatible model running locally will not always match it on every benchmark. The trade is that you can run two strong open models in parallel, compare them on your real prompts, and route the rare task that needs a particular vendor through a different path. We will tell you which side a given workload sits on.

Running on your own server adds the responsibility of being a tenant on your own infrastructure. We handle the setup, the security hardening, and the maintenance, but you should plan for a sense of who owns admin access, who approves new users, and who keeps an eye on the model choices. Most SMEs find this lighter than they expected, but it is not zero.

For how an MTAIChat setup would be sized to your team, drop us an email for more information rather than a generic figure.

MT Labs helps companies across Singapore deploy AI tools they actually own. Private infrastructure, no recurring cloud subscriptions, and a setup built around how your team already works. Whether you need a small assistant for one team or a full agentic AI for the whole company, we size the setup to what you need and what your team can manage. Get in touch and we’ll map it out with you.

FAQ

What is MTAIChat?

MTAIChat is the MT Labs self-hosted AI chat platform, built on Open WebUI and bundled on the private AI server you rent from us. The team gets a familiar chat interface, but the prompts, the chat history, and the model weights all sit on infrastructure you control. It is the private alternative to per-seat team AI chat subscriptions.

How does multi-model mode actually work?

You type one prompt and the platform sends it to two models at the same time. Both responses appear side by side in the chat. Your team reads both, keeps the better answer, blends the two, or uses one as a draft and the other as a critique. The compute happens locally on the rented server, so running a second model is hardware time, not a second subscription.

Which AI models can we run on MTAIChat?

MTAIChat works with local or OpenAI-compatible models. Practical choices include current generations of Mistral, Qwen, and Gemma, with a model builder for custom agents tuned to your work. Because it is your server, you choose which models load and you can swap them as better releases ship.

How does pricing compare with per-seat AI chat subscriptions?

The server cost is flat. Adding a user does not change the price, because MTAIChat is bundled on the rented server, not licensed per seat. Per-seat AI chat tools charge for every user every month, and adding a second model on those tools usually means a second subscription on top. Drop us an email for more information on sizing.

Can our whole company use MTAIChat, including non-technical staff?

Yes. MTAIChat supports unlimited users at a flat cost, with an admin panel for roles and granular permissions. Prefilled prompt suggestions help non-technical staff get started without a blank-box problem. The platform also runs at the MindTheory holiday camps for students, which is the practical proof that first-time users get comfortable with it quickly.

Where do our chat conversations and uploaded files actually live?

On the private AI server you rent from MT Labs. Prompts, chat history, attachments, and any documents you use for RAG sit on that machine, inside Singapore, on infrastructure you control. There is no external processor reading the messages, which removes one whole category of cross-border data transfer concern.

Can we generate images directly inside the chat?

Yes. MTAIChat is wired to the same local ComfyUI image backbone that runs SecondBrain, so image generation happens directly inside the chat from a prompt. There is no per-image cloud cost, no separate tool to open, and no usage cap beyond what the server hardware supports.

Chat with AI

Hello! I'm MTLabs AI, How can I help you today?