ローカルLLM

Making "Internal Documents Never Leave the Building" a Standard Offering — Designet Starts Selling In-House AI Infrastructure Built on Ollama, Dify, and Fess, from ¥1.4M

Making "Internal Documents Never Leave the Building" a Standard Offering — Designet Starts Selling In-House AI Infrastructure Built on Ollama, Dify, and Fess, from ¥1.4M

OSS specialist Designet launched its "In-House AI Infrastructure Construction Service" on August 26, using a local LLM to process internal documents without sending them outside. Here we lay out the RAG configuration built on Ollama, Dify, and Fess, the pricing starting from ¥1.4 million, and the bl

By FF
「社内文書は一歩も外へ出さない」を標準メニューにする──デージーネットが、Ollama・Dify・Fessで組む社内AI基盤を140万円台から売り始めた

「社内文書は一歩も外へ出さない」を標準メニューにする──デージーネットが、Ollama・Dify・Fessで組む社内AI基盤を140万円台から売り始めた

OSS専門のデージーネットが、ローカルLLMで社内文書を外に出さず処理する「社内AI基盤構築サービス」を8月26日に開始。Ollama・Dify・Fessで組むRAG構成と、140万円台からの料金、そして導入前に見ておくべき死角を整理します。

FF
The "Model Swap" Is Now a Single Toggle — Ollama Ships Official Claude Desktop Support, Letting You Run Open Models Locally for Free and Unlimited

The "Model Swap" Is Now a Single Toggle — Ollama Ships Official Claude Desktop Support, Letting You Run Open Models Locally for Free and Unlimited

Ollama ships official Claude Desktop support. With a single toggle, you can run open models (Qwen/DeepSeek/Kimi/GLM) locally or in the cloud without leaving Claude's interface. This piece lays out the practical gains in cost and privacy, how to split work between the two, and the burdens — speed, qu

By FF
「モデルの差し替え口」がトグル一つになった──OllamaがClaude Desktopに正式対応し、手元の開いたモデルを無料・無制限で回せるようにした

「モデルの差し替え口」がトグル一つになった──OllamaがClaude Desktopに正式対応し、手元の開いたモデルを無料・無制限で回せるようにした

OllamaがClaude Desktopに正式対応。トグル一つで、手元やクラウドの開いたモデル(Qwen/DeepSeek/Kimi/GLM)をClaudeの画面のまま使える。コストとプライバシーの実益、使い分けの設計、そして速度・品質・ハードウェアという負担までを整理する。

FF
"The Strongest Open Model" Shipped on Its Own Say-So — The Confident Scorecard DeepSeek V4 Pro 0813 Laid Out, and the Verification That's Still Missing

"The Strongest Open Model" Shipped on Its Own Say-So — The Confident Scorecard DeepSeek V4 Pro 0813 Laid Out, and the Verification That's Still Missing

DeepSeek released V4 Pro "0813" as GA. It touts breakthrough pricing of $0.87 per million tokens and top-tier coding scores — but most of that is self-reported, with third-party verification still pending. Here's what developers should confirm before adopting it.

By FF
「最強のオープン」は自己申告のまま出荷された──DeepSeek V4 Pro 0813が並べた強気の採点表と、まだ埋まらない裏取り

「最強のオープン」は自己申告のまま出荷された──DeepSeek V4 Pro 0813が並べた強気の採点表と、まだ埋まらない裏取り

DeepSeekがV4 Pro「0813」をGA公開。100万トークン$0.87という破格の料金とコーディング上位級のスコアを掲げるが、その大半は自己申告で第三者検証は待ち。開発者が採用前に確かめるべき点を整理します。

FF
Bringing a 1.56TB Frontier Model Down to Your Own Servers──What Kimi K3's Full Weight Release Reveals About the Cost and Fine Print of "Running It In-House"

Bringing a 1.56TB Frontier Model Down to Your Own Servers──What Kimi K3's Full Weight Release Reveals About the Cost and Fine Print of "Running It In-House"

The full weights of "Kimi K3," a largest-class open model with 2.8 trillion total parameters, were released on July 27. We lay out the reality of 1.56TB of actual data and a 64-accelerator-class cluster, the fine print of its custom license, and the one line that matters most for running it in-house

By FF
1.56TBのフロンティアモデルを自社サーバーに降ろす──Kimi K3の全重み公開が突く「内製で回す」のコストと但し書き

1.56TBのフロンティアモデルを自社サーバーに降ろす──Kimi K3の全重み公開が突く「内製で回す」のコストと但し書き

総パラメータ2.8兆の「最大級オープンモデル」Kimi K3の全重みが7月27日に公開。1.56TBの実物と64基級クラスタという現実、独自ライセンスの但し書き、そして社内利用は条件外という「内製で回す」うえで効く一文を整理します。

FF
「安さ」がモデル選びを動かした日──中国製オープンモデルが米国の開発現場を侵食し、コーディングの土台が揺れる

「安さ」がモデル選びを動かした日──中国製オープンモデルが米国の開発現場を侵食し、コーディングの土台が揺れる

7月7日、米企業の間で中国製オープンモデルの採用が急拡大していると相次ぎ報道。OpenRouterで一時46%、GLM-5.2はOpus 4.8に肉薄しコスト約5分の1。コーディングの土台となるモデル選びが、性能から「性能あたりのコスト」とデータの置き場所へ移りつつあります。

FF
手元のエージェントに「オープンな頭脳」を挿す──Moonshotの新モデルKimi K2.7-Codeが広げる選択肢

手元のエージェントに「オープンな頭脳」を挿す──Moonshotの新モデルKimi K2.7-Codeが広げる選択肢

Moonshot AI が6月12日に公開したコーディング特化モデル Kimi K2.7-Code。重みは Modified MIT で公開され、Claude Code や Cline にそのまま挿せる。開発現場の選択肢を広げる一手を、利点と留保の両面から整理します。

FF
1トークンずつ書くのをやめたGemma──Googleの「DiffusionGemma」が手元のGPUで狙う毎秒1,000トークン

1トークンずつ書くのをやめたGemma──Googleの「DiffusionGemma」が手元のGPUで狙う毎秒1,000トークン

Google DeepMindが実験的オープンモデル「DiffusionGemma」を公開。1トークンずつ書く方式をやめ、256トークンを一括生成する拡散方式で手元のGPUでも最大4倍速。Apache 2.0で誰でも入手でき、速度と品質の割り切りを開発者目線で整理します。

FF