DeepSeek Harness, MiniMax Music 3, and Gemini 3.7 Flash Open‑Source: Architecture and Benchmarks

The article announces the open‑source release of DeepSeek Harness with a plugin‑centric architecture and four operational modes, introduces MiniMax Music 3 capable of generating five‑minute songs using dual language models, and details Gemini 3.7 Flash’s performance gains across coding, web‑UI, and knowledge‑intensive benchmarks while highlighting its competitive pricing.

SuanNi
SuanNi
SuanNi
DeepSeek Harness, MiniMax Music 3, and Gemini 3.7 Flash Open‑Source: Architecture and Benchmarks

DeepSeek Harness has been released under the MIT license as a developer preview (v0.1). It follows an "everything‑as‑plugin" design driven by Cordis, inspired by the paper *A Programming Paradigm for Spatiotemporal Composability*. The system logs all interactions in an append‑only session log, including system prompts, chain‑of‑thought, tool calls, sub‑agent scheduling, and context injections, which can be inspected in the Trajectory view for recovery, forking, and replay.

The platform offers four preset modes, each loading a different plugin set: Standard (full tool suite), PTC (Programmatic Tool Calling, where model‑generated code orchestrates multi‑turn tool usage), Minimal (only a shell tool and a file‑edit tool for lightweight benchmarking), and Creative (allows runtime inspection and composition of Cordis plugins to craft new modes).

MiniMax Music 3, the next‑generation production‑grade music model, supports generating complete five‑minute songs that maintain thematic, rhythmic, vocal, and melodic continuity. It combines an 800‑million‑parameter global language model for long‑range musical structure with a 60‑million‑parameter local model for frame‑level audio detail. The model employs a continuous hidden‑state synthesis system based on flow‑matching and explicit Bayesian algorithms, outputting 32 kHz, 16‑bit stereo WAV files.

Gemini 3.7 Flash, Google’s latest model for programming and intelligent agents, shows notable improvements in software engineering, knowledge work, and web‑development workflows. In coding tasks, FrontierCode 1.1 Main achieves 43.6% accuracy, while DeepSWE v1.1 reaches 65.3%, both surpassing the previous 3.6 Flash. In web‑UI generation, Arena.ai’s WebDev Arena test gives Gemini 3.7 Flash an Elo score of 1588, ranking 8th globally. For knowledge‑intensive domains (finance, law, biosciences), the GDP.pdf benchmark reports a 34.0% score increase over 3.6 Flash. AutomationBench scores rise to 30.4% from 17.0%.

Despite its superior performance, Gemini 3.7 Flash remains cost‑effective, and its pricing continues to undercut competitors, contributing to an escalating large‑model price war alongside GPT 5.6 Terra’s price cut and Meta’s Muse Spark 1.2 discount.

Reference links include DeepSeek Harness’s GitHub repository, MiniMax Music 3’s GitHub and Hugging Face pages, and the official announcements on X for each project.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

AI agentslarge language modelsBenchmarkDeepSeek HarnessGemini 3.7 FlashMiniMax Music 3
SuanNi
Written by

SuanNi

A community for AI developers that aggregates large-model development services, models, and compute power.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.