Tagged articles

Gemma

16 articles · Page 1 of 1
macrozheng
macrozheng
Sep 23, 2026 · Mobile Development

Run LLMs on 6GB Android Phones with OlliteRT: OpenAI-Compatible LAN API

OlliteRT is an open-source Android app that runs large language models locally via Google's LiteRT runtime, exposing an OpenAI-compatible API over LAN for lightweight always-on tasks like RSS summarization and smart home automation, with monitoring, logging, and security features.

AndroidEdge AIGemma
0 likes · 11 min read
Run LLMs on 6GB Android Phones with OlliteRT: OpenAI-Compatible LAN API
Architecture Digest
Architecture Digest
Sep 2, 2026 · Artificial Intelligence

OlliteRT: Turn Old Android Phones into Local LLM API Servers

OlliteRT is an open-source Android app that transforms old phones into local LLM servers exposing OpenAI-compatible APIs, supporting models like Gemma 4 E2B and Qwen 2.5 1.5B with configurable hardware acceleration, monitoring, and Home Assistant integration, all running offline on 5-10W power.

AndroidEdge AIGemma
0 likes · 10 min read
OlliteRT: Turn Old Android Phones into Local LLM API Servers
AI Engineering
AI Engineering
Sep 1, 2026 · Artificial Intelligence

Why Long-Term LLM Agents Fail: The Same Design Choice Behind Two Deaths

Long‑term LLM agents suffer from ever‑slowing execution and context poisoning because they continuously append every observation, action, and reasoning step to the prompt, but the SKILL.state approach replaces this growing history with a compact mutable state, dramatically cutting token usage while boosting accuracy and robustness across diverse benchmarks.

BenchmarkGeminiGemma
0 likes · 11 min read
Why Long-Term LLM Agents Fail: The Same Design Choice Behind Two Deaths
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 29, 2026 · Artificial Intelligence

Can Chinese‑Only Inference Training Match English? Apple’s New Study Shows Only 1.1% Gap

Apple and the Hasso Plattner Institute evaluated over 200 multilingual inference training experiments across nine base models and eleven languages, finding that training with Chinese rewards incurs just a 1.1 percentage‑point loss versus English, while low‑resource languages can cause severe performance collapses in specific model‑language combos.

GemmaQwen3cross-language transfer
0 likes · 9 min read
Can Chinese‑Only Inference Training Match English? Apple’s New Study Shows Only 1.1% Gap
Java Tech Enthusiast
Java Tech Enthusiast
Aug 14, 2026 · Artificial Intelligence

Will Codex Become Obsolete? Insights from HuggingFace Harness Experiments

The article examines recent HuggingFace experiments comparing how different harnesses affect large and small AI models, revealing that complex harnesses like Codex excel on big models while lightweight harnesses perform better on smaller ones, and discusses why future agents will demand far more compute than current setups.

AIAgentGLM
0 likes · 6 min read
Will Codex Become Obsolete? Insights from HuggingFace Harness Experiments
HyperAI Super Neural
HyperAI Super Neural
Jun 12, 2026 · Artificial Intelligence

DiffusionGemma Boosts Text Generation Speed Up to 4× with Discrete Diffusion

Google’s open‑source DiffusionGemma model leverages a 26‑billion‑parameter Mixture‑of‑Experts architecture and discrete diffusion decoding to generate whole text blocks, achieving up to four times faster generation—over 1100 tokens/s on an NVIDIA H100 and 700 tokens/s on an RTX 5090—while activating only 3.8 billion parameters during inference.

DiffusionGemmaDiscrete DiffusionGPU acceleration
0 likes · 4 min read
DiffusionGemma Boosts Text Generation Speed Up to 4× with Discrete Diffusion
Geek Labs
Geek Labs
Apr 11, 2026 · Mobile Development

PhoneClaw and PokeClaw: Turning Your Phone into a Private AI Agent

PhoneClaw and PokeClaw are open‑source, on‑device AI agents for iOS and Android that run Gemma 4 locally, offering offline privacy, zero‑cost operation, and native tool calling through iOS APIs or Android Accessibility services.

AndroidGemmaPhoneClaw
0 likes · 11 min read
PhoneClaw and PokeClaw: Turning Your Phone into a Private AI Agent
AI Explorer
AI Explorer
Apr 8, 2026 · Artificial Intelligence

Exploring Google AI Edge Gallery: Running Large Models Locally on Your Phone

Google’s AI Edge Gallery lets you run cutting‑edge large language models such as Gemma 4 entirely offline on Android or iOS devices, offering absolute privacy, zero‑latency responses, and a modular platform with agent skills, thinking mode, multimodal input, and a prompt‑lab for on‑device AI experimentation.

GemmaGoogle AI Edge GalleryKotlin
0 likes · 6 min read
Exploring Google AI Edge Gallery: Running Large Models Locally on Your Phone
AI Explorer
AI Explorer
Apr 7, 2026 · Mobile Development

Google AI Edge Gallery: Offline Mobile AI with Gemma Models and Multimodal Agents

Google’s AI Edge Gallery lets developers run open‑source large language models such as Gemma 4 directly on Android devices without network connectivity, offering an integrated framework with agent skills, thinking mode visualizations, multimodal interaction, and a prompt lab, thereby addressing privacy, latency, and offline AI needs.

AndroidGemmaGoogle AI Edge Gallery
0 likes · 6 min read
Google AI Edge Gallery: Offline Mobile AI with Gemma Models and Multimodal Agents
Code Mala Tang
Code Mala Tang
Jul 22, 2025 · Artificial Intelligence

Convert Any PDF to Clean Markdown with a Local LLM (Gemma 3)

Learn how to transform any PDF—including scanned documents—into well‑structured Markdown using a local LLM (Gemma 3 via Ollama), Python, PyMuPDF and Pillow, without cloud APIs or API keys, by converting pages to images, prompting the model, and saving the output.

GemmaLLMOllama
0 likes · 12 min read
Convert Any PDF to Clean Markdown with a Local LLM (Gemma 3)
Rare Earth Juejin Tech Community
Rare Earth Juejin Tech Community
Feb 23, 2024 · Artificial Intelligence

Google’s Open‑Source Gemma Large Language Model: Architecture, Performance, and Community Reception

Google has released the open‑source Gemma LLM series (2B and 7B parameters) built on Gemini‑style architecture, offering free, commercial‑ready models that run on notebooks, support JAX/PyTorch/TensorFlow, outperform many open‑source peers, and have quickly sparked extensive community testing and discussion.

Artificial IntelligenceGemmaGoogle
0 likes · 5 min read
Google’s Open‑Source Gemma Large Language Model: Architecture, Performance, and Community Reception
21CTO
21CTO
Feb 22, 2024 · Artificial Intelligence

How Google’s Open‑Source Gemma Model Brings LLM Power to Your Laptop

Google’s newly released open‑source Gemma models let developers run powerful large‑language‑model workloads on notebooks, workstations, or cloud platforms, offering competitive performance, extensive tooling, and built‑in safety measures for responsible AI deployment.

AI safetyGemmaGoogle AI
0 likes · 6 min read
How Google’s Open‑Source Gemma Model Brings LLM Power to Your Laptop
Programmer DD
Programmer DD
Feb 22, 2024 · Artificial Intelligence

Google Unveils Gemma: Open‑Source LLM Matching Gemini’s Power

Google has launched Gemma, an open‑source large language model available in 2B and 7B parameter versions, built on the same technology as Gemini, outperforming many existing models and capable of running on ordinary laptops, with a detailed technical report and quick‑start guide provided online.

AIGemmaGoogle
0 likes · 3 min read
Google Unveils Gemma: Open‑Source LLM Matching Gemini’s Power