作者: hiyoho

  • Voicebox:本地优先的开源 AI 语音工作室,克隆声音 + 让智能体开口说话

    Voicebox:本地优先的开源 AI 语音工作室,克隆声音 + 让智能体开口说话

    Voicebox 开源 AI 语音工作室

    Voicebox —— 本地优先的开源 AI 语音工作室

    📌 项目简介

    Voicebox 是由 Jamie Pine 打造的开源 AI 语音工作室,一句话概括:克隆任意声音、生成语音、随时听写、还能让你的 AI 智能体用你拥有的声音开口说话。它是 ElevenLabs(语音输出)与 WisprFlow(语音输入)的本地优先、免费开源二合一替代品,整套语音 I/O 栈都跑在你自己的机器上。

    ⭐ GitHub 43,196 stars | 🍴 5,280 forks | 📅 创建于 2026-01-25,更新于 2026-07-19(当日 GitHub Trending 热门)| 📜 MIT 许可 | 💻 TypeScript / Tauri(Rust) / FastAPI(Python)

    🛠 安装要求和过程

    环境要求

    • 普通用户:直接下载桌面应用,无需配置开发环境;支持 macOS(Apple Silicon / Intel)、Windows、Linux(需源码构建)、Docker。
    • GPU 加速:macOS 走 MLX/Metal(神经引擎快 4-5×),Windows/Linux NVIDIA 走 CUDA,AMD 走 ROCm,Intel Arc 走 IPEX/XPU,Windows 任意显卡走 DirectML,无显卡则回退 CPU。
    • 开发者构建:需 BunRustPython 3.11+Tauri 前置依赖(macOS 还需 Xcode)。

    快速安装(三种方式)

    方式一 · 下载安装包(推荐)

    方式二 · 开发者一键启动

    git clone https://github.com/jamiepine/voicebox.git
    cd voicebox
    just setup    # 创建 Python venv 并安装全部依赖
    just dev      # 启动后端 + 桌面应用

    (先安装 just 命令运行器:brew install justcargo install just

    方式三 · 接入 AI 智能体(MCP)

    claude mcp add voicebox \
      --transport http \
      --url http://127.0.0.1:17493/mcp \
      --header "X-Voicebox-Client-Id: claude-code"

    ✨ 核心功能

    Voicebox 界面

    1. 7 大 TTS 引擎,自由切换:Qwen3-TTS、Qwen CustomVoice、LuxTTS、Chatterbox Multilingual、Chatterbox Turbo、HumeAI TADA、Kokoro,覆盖 23 种语言(英/阿/日/印地/斯瓦希里等),并支持零样本声音克隆与 50+ 预设音色。
    2. 声音克隆 + 语音性格:几秒参考音频即可克隆任意声音;可为每个声纹绑定自由格式的”性格”,通过本地 Qwen3 LLM 实现 Compose(即兴台词)与 Speak-in-character(拟人改写)。
    3. 全局听写 / 语音输入:系统级热键按住说话、松手即停(push-to-talk),macOS 下自动粘贴到当前输入框;底层基于 OpenAI Whisper 的 STT,支持 Base/Small/Medium/Large/Turbo。
    4. 后期音效处理:基于 Spotify pedalboard 的 8 种效果——变调、混响、延迟、合唱、压缩、增益、高/低通滤波,可实时预览并存为预设。
    5. Agent 语音输出(MCP):一句 voicebox.speak() 调用,任何支持 MCP 的智能体(Claude Code / Cursor / Cline / Windsurf)都能用你克隆的声音向你播报任务进展。

    Voicebox Stories 编辑器

    🎯 典型使用场景

    • 播客 / 有声内容创作:用内置 Stories 多轨时间线编辑器,零样本克隆嘉宾声音,一键生成多角色对话、播客与叙事音频,长文本自动分块 + 交叉淡入淡出。
    • AI 编码智能体开发循环:把 Claude Code 绑定到某个克隆声纹,构建完成、测试通过时它能”开口”告诉你——用耳朵接收进度,而不是一直盯着终端。
    • 无障碍与语音辅助:为无法用原声说话的人提供克隆声音;全局热键听写 + macOS 自动粘贴,大幅提升文字录入效率。
    • 游戏 / 叙事交互:为游戏角色或互动叙事工具接入带”性格”的语音,实现有情绪、带语气词的拟人对话。

    💡 推荐理由

    我用过多家云端语音服务,Voicebox 最打动我的是三点:

    • 真正的本地优先 + 隐私优先。模型、声音数据、录音全部留在本机,不上传任何云,对注重数据主权的用户极其友好。
    • 输入 + 输出一站式。ElevenLabs 只管”说”,WisprFlow 只管”听”,Voicebox 把两者打通,还用一个本地 LLM 串联性格改写,一 GPU 显存占用搞定全套。
    • 给 AI Agent 配声音这个设计非常新颖。当你的编码助手能用你熟悉的声音播报”部署完成”,人机交互的体验立刻不一样,这也是它和同类工具最大的差异化。

    ⚠️ 一点不足:Linux 目前没有预编译二进制,需要源码构建;同时 7 个引擎质量参差,部分小模型音色一般,建议按场景挑选。但瑕不掩瑜,作为开源语音 I/O 基础设施,它已经非常能打。

    📥 下载地址

    本文由自动化脚本基于 GitHub 公开仓库信息整理,项目数据截至 2026-07-19。

  • 26岁东亚美女黑裙毛皮披肩人像

    26岁东亚美女黑裙毛皮披肩人像

    26岁东亚美女黑裙毛皮披肩人像



    🤖 ChatGPT

    🇺🇸 English Prompt

    Approximately 26-year-old East Asian beauty, black hair neatly tied into a high bun, natural side-parted bangs. Exquisite face, bright red lipstick, confident and gentle smile. Realistic skin texture and natural matte glow.
    
    Black off-shoulder dress paired with a fluffy black fur shawl, revealing the collarbone. Exquisite silver necklace and small earrings. Black lace floral pattern stockings. Sitting on a bright yellow modern chair, holding a gold chain bag in her right hand, natural selfie or casual portrait angle.
    
    Realistic smartphone photo style shot by iPhone 16 Pro, natural indoor lighting, soft top and side lighting. Modern dark wood wall background, shallow depth of field. Hyper-realism, realistic photo texture, natural colors, meticulous details --ar 9:16

    🇨🇳 中文提示词

    大约26岁的东亚系美女,黑发整齐地束成高高的发髻,自然的侧分刘海。精致的面容,明艳的红色唇膏,充满自信的温柔微笑。逼真的皮肤质感和自然的哑光光泽。
    
    黑色露肩连衣裙搭配蓬松的黑色毛皮披肩,露出锁骨。精致的银色项链和小巧的耳环。黑色蕾丝花卉图案丝袜。坐在明亮的黄色现代椅子上,右手拿着金色链条包,自然的自拍或休闲肖像角度。
    
    iPhone 16 Pro拍摄的逼真智能手机照片风格,自然的室内照明,柔和的顶光和侧光。现代深色木墙背景,浅景深。超现实主义,逼真的照片质感,自然的色彩,细致的细节 --ar 9:16
  • 超写实清晨刚醒亲密自拍

    超写实清晨刚醒亲密自拍

    超写实清晨刚醒亲密自拍



    🤖 ChatGPT

    🇺🇸 English Prompt

    Create a high-quality ultra-realistic intimate morning selfie of the same adult person immediately after waking up.
    The image should feel like a spontaneous real-life moment casually captured with a smartphone front-facing camera.
    Scene and pose
    The subject is lying face-down on a bed, with their back facing the window and their face turned toward the camera.
    The left side of the face is deeply pressed into a soft pillow, naturally compressing the cheek.
    The camera is extremely close to the face, creating an intimate front-camera selfie perspective.
    Only one person should appear.
    Show mainly the face and a small amount of the upper body.
    The subject should look:
    sleepy
    slightly dazed
    quiet
    naturally soft
    as if they have only just opened their eyes
    The eyes should remain relaxed and slightly unfocused.
    Use a calm, absent-minded gaze toward the camera.
    Preserve the original hairstyle and hair color from the uploaded portrait.
    Make the hair extremely messy and naturally tousled from sleep.
    Include:
    loose uneven strands
    soft bed-head volume
    slightly flattened areas from the pillow
    realistic flyaway hairs
    The hair should look genuinely unstyled after waking up, not intentionally styled.
    Makeup
    Apply refined but natural Korean beauty makeup while preserving the original facial identity.
    luminous natural skin
    subtle soft blush
    lightly defined eyebrows
    delicate neutral eye makeup
    thin eyeliner following the original eye shape
    softly separated lashes
    glossy natural pink lips
    The makeup should remain polished but believable.
    Do not heavily reshape or beautify the face.
    The subject wears a fitted white long-sleeve top.
    Only a small portion of the upper clothing should be visible because the composition is focused closely on the face.
    Lighting and background
    Use dark black-gray curtains behind the subject.
    Morning light enters only through a narrow gap in the curtains from behind.
    Create strong but realistic contrast between light and shadow across the face.
    Keep the face clearly visible and softly illuminated despite the backlighting.
    Use subtle morning glow, natural shadow transitions, and realistic indoor light.
    Composition and camera
    extreme facial close-up
    smartphone front-camera perspective
    csual imperfect framing
    intimate eye-level angle
    shallow depth of field
    realistic lens softness
    no professional studio posing
    The final result should feel like an authentic private morning selfie captured without preparation.
    Image quality
    Create premium ultra-realistic 4K-quality detail with:
    natural high-quality skin texture
    visible pores
    realistic individual hair strands
    soft pillow texture
    natural facial compression against the pillow
    accurate anatomy
    realistic morning lighting
    Final mood:
    A sleepy, intimate, freshly awakened morning selfie that feels like a genuine everyday moment.

    🇨🇳 中文提示词

    创建一张高质量超写实的亲密清晨自拍,同一位成年人刚睡醒后的样子。
    这张照片应该感觉像是一个自发的现实生活瞬间,用智能手机前置摄像头偶然捕捉到的。
    场景和姿势
    主体脸朝下趴在床上,背对着窗户,脸转向镜头。
    脸的左侧深深地压在柔软的枕头上,自然地挤压着脸颊。
    相机离脸非常近,营造出一种亲密的前置摄像头自拍视角。
    只应该出现一个人。
    主要展示面部和少量的上半身。
    主体看起来应该是:
    困倦的
    略显迷茫
    安静
    自然柔和
    仿佛刚刚睁开眼睛
    眼睛应该保持放松且略微失焦。
    向镜头投去冷静、心不在焉的目光。
    保留上传肖像中的原始发型和发色。
    让头发极其凌乱,带有自然的睡眠后的蓬乱感。
    包括:
    散乱不均的发丝
    柔软的睡醒后的蓬松感
    被枕头压平的局部区域
    真实的碎发
    头发看起来应该是真正的刚醒后未打理的样子,而不是刻意造型过的。
    妆容
    应用精致但自然的韩式美容妆容,同时保留原始的面部特征。
    发光的自然肌肤
    微妙柔和的腮红
    浅淡定义的眉毛
    精致的中性眼妆
    顺着原始眼形的细眼线
    根根分明的睫毛
    光亮的自然粉色嘴唇
    妆容应保持精致但可信。
    不要过度重塑或美化面部。
    主体穿着紧身的白色长袖上衣。
    由于构图紧贴面部,只能看到一小部分上身衣物。
    照明和背景
    主体身后使用深黑灰色的窗帘。
    晨光仅通过身后窗帘的狭窄缝隙进入。
    在脸上创造出强烈但真实的明暗对比。
    尽管有背光,但要保持面部清晰可见且光线柔和。
    使用微妙的晨光、自然的阴影过渡和真实的室内光线。
    构图和相机
    极端面部特写
    智能手机前置摄像头视角
    随意不完美的构图
    亲密的平视角度
    浅景深
    真实的镜头柔和感
    没有专业的摄影棚摆拍
    最终结果应该感觉像是一张真实、私密、未经准备捕捉到的清晨自拍。
    图像质量
    创建具有以下特征的高级超写实4K画质细节:
    高质量的自然皮肤纹理
    可见的毛孔
    真实的单根发丝
    柔软的枕头纹理
    自然的靠在枕头上的脸部挤压感
    准确的解剖结构
    真实的清晨光效
    最终氛围:
    一个困倦、亲密、刚睡醒的清晨自拍,感觉像是一个真实的日常生活瞬间。
  • 简约室内粉色花朵裙少女写实肖像

    简约室内粉色花朵裙少女写实肖像

    简约室内粉色花朵裙少女写实肖像



    🤖 ChatGPT

    🇺🇸 English Prompt

    Ultra-realistic portrait of a young woman against a plain light gray wall with a clean minimalist indoor background. She wears a pastel blush pink floral chiffon halter dress featuring watercolor flower prints, sheer off-shoulder puff sleeves, ruffled trim, and a handmade chiffon flower at the neckline. Her long dark brown hair is styled in a thick side braid with soft waves, airy curtain bangs, and loose face-framing curls. She poses naturally with shoulders slightly angled, head gently tilted, eyes gazing off-camera, and lips slightly parted, creating a soft, dreamy, feminine expression. Makeup: flawless porcelain glass skin, satin foundation, seamless concealer, soft jawline and nose contour, peach-pink blush blended across the cheeks and nose bridge, champagne highlighter on cheekbones, nose, brow bone, cupid’s bow, and inner corners, soft brown brows with natural hair strokes, muted pink and taupe gradient eyeshadow, subtle shimmer lids, delicate aegyo-sal, thin brown winged eyeliner, long wispy curled lashes, defined lower lashes, glossy rosy gradient lips, luminous neck, shoulders, and collarbones. Eye-level medium close-up, Lighting: aggressive direct sunlight + on-camera flash. Underexposed -1.5 EV for dark luxury mood. add small aesthetic signature watermark sabine. Shot on Canon PowerShot G7X Mark III, pop flash, sharp eyes, 35mm film grain. --ar 4:5 -v.6.1 do not alter face

    🇨🇳 中文提示词

    一位年轻女性的超写实肖像,背景是简洁的浅灰色墙壁和干净的极简室内环境。她穿着一件淡粉色花朵雪纺挂脖连衣裙,上面有水彩花卉印花,带有透明的露肩泡泡袖、荷叶边装饰,领口处有一朵手工雪纺花。她的深褐色长发被编成厚实的侧边辫子,带有柔和的波浪、轻盈的法式刘海和修饰脸型的碎发。她姿势自然,肩膀微侧,头轻轻倾斜,目光望向镜头外,双唇微张,营造出柔和、梦幻且充满女性气质的表现。妆容:完美无瑕的瓷感玻璃肌、缎面底妆、无痕遮瑕、柔和的下颌线和鼻梁修容、晕染在双颊和鼻梁上的桃粉色腮红、在颧骨、鼻子、眉骨、唇峰和内眼角扫上的香槟色高光、带有自然毛流感的柔和棕色眉毛、哑粉色和灰褐色渐变眼影、微妙闪光的眼睑、精致的卧蚕、细长的棕色猫眼眼线、纤长卷曲的睫毛、根根分明的下睫毛、水光的玫瑰色渐变唇、有光泽的脖颈、肩膀和锁骨。平视中景特写,灯光:强烈直射阳光 + 相机闪光灯。曝光补偿 -1.5 EV 以营造深色奢华氛围。添加小型的美学签名水印 sabine。由佳能 PowerShot G7X Mark III 拍摄,弹出式闪光灯,犀利的眼神,35mm 胶片颗粒。--ar 4:5 -v.6.1 请勿改变面部。
  • 多乐士油漆液体雕塑探戈舞者艺术海报

    多乐士油漆液体雕塑探戈舞者艺术海报

    多乐士油漆液体雕塑探戈舞者艺术海报



    🤖 ChatGPT

    🇨🇳 中文提示词

    Create an ultra-premium editorial advertising poster for Dulux, centered on a hyper-real female tango dancer whose dress and sweeping motion trails are formed entirely from sculptural liquid paint rising directly from a Dulux paint can below. The image must feel bold, elegant, dramatic, fashion-forward, and globally premium, with a strong high-contrast pure-color background and a powerful sense of rhythm, tension, and movement. The final result should look like a world-class Dulux campaign poster where color becomes passion, paint becomes couture, and dance becomes visual impact.
    
    Core concept:
    The hero visual is a refined young Western female dancer captured in a rhythmic tango pose full of control, confidence, and mature dramatic energy. Her liquid-paint costume rises directly from a Dulux paint can below and forms both the sculptural body of the dress and the sharp sweeping skirt extensions associated with tango motion. The visual language should feel more sensual, grounded, and directional than ballet: stronger hip line, firmer posture, sharper turning energy, and a more theatrical silhouette. The image must feel iconic, modern, and unforgettable.
    
    Composition:
    Use a vertical luxury poster layout with the dancer placed centrally and slightly above the lower third, creating a commanding, highly readable silhouette. The Dulux paint can should sit near the lower center or slightly offset near the grounded foot, clearly visible as the source of the paint. The liquid paint must rise from the can in a continuous elegant stream, forming the fitted upper dress and dramatic tango-like skirt sweeps around the legs. The composition should feel clean, graphic, and intense, with strong negative space so the gesture and color contrast dominate.
    
    Female dancer:
    Use a refined young Western female model with fair skin, elegant facial structure, healthy graceful physique, and a poised fashion-editorial presence. Her pose should suggest tango rhythm and emotional charge:
    strong grounded supporting leg,
    one leg extending or crossing with precision,
    twisted torso,
    sharp arm line or lifted hand,
    elongated neck,
    and a confident, composed expression.
    She must feel powerful, feminine, premium, and visually magnetic.
    
    Liquid paint tango dress:
    The dress must be formed entirely from hyper-real liquid paint. The paint should create a fitted bodice and a dramatic, asymmetrical tango skirt, with sweeping arcs and directional trails that feel like fabric in motion but remain unmistakably liquid. The paint must be glossy, rich, thick, and physically convincing, with refined edges, controlled splash tension, and elegant sculptural weight. The movement should feel passionate and rhythmic, not chaotic.
    
    Paint source logic:
    The Dulux paint can must clearly act as the source of the entire liquid construction. The viewer should instantly understand that the paint rises from the can and transforms into the tango dress and dynamic motion trails. This physical connection is essential.
    
    Background:
    Use a bold high-saturation pure-color background with no clutter and no scenery. For the strongest tango tension, use a vivid deep teal or saturated emerald-teal background contrasted against a rich crimson-red or lacquered scarlet paint dress. The background should feel flat, clean, modern, and poster-like, intensifying the silhouette and making the liquid dress feel more explosive and luxurious.
    
    Color strategy:
    The paint dress should feel jewel-like, deep, and sensual, using crimson red, lacquer red, or deep vermilion as the main liquid color. The background should contrast strongly through a cool saturated tone such as deep teal or peacock green. Keep the dancer’s skin tones natural and refined so the main chromatic energy comes from the red-versus-teal contrast. The final image should feel bold, artistic, expensive, and globally branded.
    
    Product integration:
    Include a hyper-real Dulux paint can or bucket with clear logo visibility in the lower area of the composition. The can should be sharply rendered, premium, and proportionally balanced within the overall poster.
    
    Typography:
    Use an elegant minimal English-led typography system with a refined Chinese subtitle.
    Suggested text direction:
    Main English title:
    “COLOUR IN RHYTHM”
    or
    “TANGO IN COLOUR”
    Chinese subtitle:
    “色彩探戈”
    or
    “让色彩起舞”
    Supporting English copy:
    “Dulux turns rhythm into colour.”
    Typography should remain minimal, modern, artistic, and beautifully spaced, allowing the figure, paint, and contrast to dominate.
    
    Brand coding:
    Preserve Dulux identity through:
    premium paint can visibility,
    confident use of bold color,
    clean international advertising polish,
    minimal composition,
    and the direct transformation of paint into fashion and movement.
    The overall result must feel like a refined same-series Dulux luxury campaign with stronger drama and emotional punch.
    
    Lighting:
    Use soft premium studio lighting with refined highlights on the dancer’s skin, hair, and liquid paint dress. The red paint must show glossy reflections, sculptural depth, and elegant contour shaping. The figure should remain dimensional and luxuriously lit against the saturated teal background, with no harsh shadow clutter.
    
    Mood:
    bold, dramatic, elegant, feminine, passionate, premium, fluid, fashion-forward, high-contrast, striking, globally branded, minimalist
    
    Rendering style:
    hyper-real Dulux advertising poster, female tango dancer formed with liquid paint dress, saturated contrasting pure-color background, luxury editorial typography, premium studio lighting, international print campaign aesthetic, 8k, world-class image quality
    
    Negative prompt:
    messy paint splash, cartoon dancer, cluttered background, weak Dulux branding, low-detail liquid, cheap fashion ad, rough typography, unrealistic anatomy, muddy colors, weak contrast, low-end poster design
  • 高密度有机蔬菜平面视觉海报

    高密度有机蔬菜平面视觉海报

    高密度有机蔬菜平面视觉海报



    🤖 ChatGPT

    🇺🇸 English Prompt

    Establish a high-density flat visual around any subject object, first refining the subject-related forms into flat organic silhouettes of vastly different sizes and varying edges, letting rounded and wide large outlines, slender lines, fine segmented forms, and scattered small pieces reach in from all sides and cross the edges, forming a dense coverage across the entire frame with no obvious perspective. Outlines are layered step by step with transparent ink, the same form repeatedly scaled, rotated, and obscured, naturally generating new color blocks where deep and light shades meet; at the same time, intentionally preserve the irregular bright base color running through them, cutting the most key subject symbols directly from these blank spaces, using only minimalist outer outlines and a few identifying details, making them flicker in and out of view, requiring a second look to be discovered. The blank images and colored silhouettes must share boundaries and intersperse with each other, and cannot become independent icons pasted on the surface. The whole is completed in a way that combines flat paper collage with transparent water-based overprinting: large shape edges are clear but with slight manual cutting irregularities, partially retaining real material textures, fiber feel, fine veins, slight ink layer thickness variations, and overprint color differences, without adding three-dimensional light and shadow, glass texture, or uniform aging. Colors are derived from the emotions, semantics, and material associations of future themes, choosing a dominant hue to form a high-coverage medium-brightness color group, with deeper homologous colors bearing the occlusion and visual weight, lighter low-saturation colors bearing the breathability and paper feel, and a few bright areas close to the base color responsible for separating shapes; controlled within a narrow color gamut, deep colors are full but not dull black, light colors are clean and not overexposed, remaining refreshing, natural, quiet, and full of life after overprinting. The layout adopts asymmetrical balance: graphics fill the screen and are cut at the edges, a readable information area is left on one side below, the subject title is arranged into three to four lines of compact white text, using simple narrow sans-serif and slightly wider East Asian character skeletons, with fine horizontal lines of varying lengths and slight manual skew interspersed between lines; a small foreign language subtitle is placed above the title, and time, place, and description are organized below with a sharp drop in font size, letter spacing is slightly open, and white text is pressed directly onto the dark collage. A small geometric logo is placed on the upper edge of the other side, and extremely small author or column text is embedded in the upper-middle part of the picture, letting information be surrounded by organic forms yet remaining clear. The reading order is first attracted by the full-frame overlapping colors, then discovering the blank symbols, and finally falling to the title below; always maintaining the tension of density, transparency, concealment, and manifestation, avoiding arranging silhouettes into regular patterns and avoiding letting text steal the surprise of negative shape discovery. Theme: Summer garden fruits and vegetables, organic vegetables, poster.

    🇨🇳 中文提示词

    围绕任意主题对象建立高密度平面视觉,先把主题相关形态提炼为大小悬殊、边缘各异的扁平有机剪影,让圆润宽阔的大块轮廓、细长线条、细密分节形态和零散小片从四周伸入并越过边缘,在整个画面形成无明显透视的密集覆盖。轮廓以透明墨色逐层压叠,同一形态反复缩放、旋转、遮挡,深浅交汇处自然生成新的色块;同时有意保留贯穿其间的不规则明亮底色,把最关键的主题符号直接从这些空白中剪出,只用极简外轮廓与少量识别细节,使其忽隐忽现、需要二次观看才被发现。空白形象与有色剪影必须共用边界并彼此穿插,不能变成贴在表面的独立图标。整体以平面纸艺拼贴结合透明水性套印的方式完成:大块形状边缘清楚但略带手工切割的不匀,局部保留真实素材纹理、纤维感、细脉线、轻微墨层浓淡和叠印色差,不添加立体光影、玻璃质感或均匀旧化。色彩从未来主题的情绪、语义与物质联想中派生,选一个主导色相构成高覆盖的中明度色群,以较深同系色承担遮挡与视觉重量、较浅低饱和色承担透气与纸面感,少量接近底色的明亮区域负责分离形状;控制在窄色域内,深色饱满但不闷黑,浅色洁净不过曝,叠印后仍保持清爽、自然、安静而富有生命感。版面采用不对称平衡:图形铺满并在边缘截断,下方一侧留出可读信息区,将主题标题排成三至四行紧凑白字,使用朴素窄体无衬线与略宽的东亚文字骨架,行间穿插长度不一、略带手工偏斜的细横线;标题上方放小号外文副题,下方按字号骤降组织时间、地点和说明,字距稍开,白字直接压在深色拼贴上。另一侧上缘放小型几何字标,画面中上部嵌入极小的作者或栏目文字,让信息像被有机形态包围却仍保持清晰。阅读顺序先被满幅叠色吸引,再发现空白符号,最后落到下方标题;始终保持密、透、藏、显的张力,避免把剪影排成规则纹样,也避免让文字抢走负形发现的惊喜。
    
    主题:夏日菜园蔬果 有机蔬菜 海报
  • 上线四天就认怂:Meta 用 @ 别人就能造 AI 深伪图,被骂到下架

    一个 @ 就能把陌生人塞进你的图里

    这周二,Meta 高调推出了自家的图像生成模型 Muse Image,并把它接进了 Instagram。其中一个被当成亮点的玩法很直白:在 Meta AI 里 @ 某个公共 Instagram 账号,系统就会把那个人的内容拉进你生成的图里。

    问题出在默认设置上。只要账号是公开的,它的照片就能被别人拿去生成 AI 图像,而且不需要本人同意。想不被用?得自己钻进设置里一层层翻,手动关掉。

    追求高风险的设计,却把责任推给个人去层层跳转菜单退出——这是不可接受的。

    说这话的是 Haley McNamara,美国全国性反性剥削组织(NCOSE)的负责人。她的担忧很具体:这不光是侵蚀了每个人对自己肖像的权利,更像是一个为”色情勒索”和其他诈骗者准备好的现成工具。

    四天,从上线到撤回

    批评声不只是来自民间。演员工会 SAG-AFTRA 直接建议旗下会员去关掉这个功能,还手把手写了操作教程。在舆论、行业组织和媒体一起施压之后,Meta 认了。

    公司在博客更新里写:”我们听到了反馈,这个功能没达到预期,所以不再提供了。”不过要注意,被撤下的只是”@ 公共账号生成图像”这一项,Muse Image 本身还能在 Meta AI、Instagram Stories 和 WhatsApp 里继续用。

    Meta Muse Image 的 @ 提及功能
    用 @ 就能把公开账号的内容拉进 AI 生成图(图:The Verge)

    真正的争论在”同意”两个字

    这起翻车背后,是一场关于”Opt-in 还是 Opt-out”的拉锯。Meta 选了后者:默认你同意,不想被用就自己去关。可当用户对自己形象的控制权,取决于他是否碰巧发现了那个隐藏开关,这套逻辑本身就站不住脚。

    • 演员公会主张,数字分身这类事必须”显式同意”,而不是事后退出;
    • 反剥削组织指出,把伦理风险从设计公司转嫁到使用者头上,是硅谷反复出现的老毛病;
    • 而随着 AI 图像、视频越来越难用肉眼分辨,隐形水印和内容标签也只是杯水车薪。

    有意思的是时间线。从官宣到撤回,前后不过四天。放在几年前,一个功能引发争议、公司再慢慢评估,往往是以”周”甚至”月”计的。2026 年的舆论场,反应速度快到可以用天来量。Meta 砸重金重建 AI 实验室、和 OpenAI、Anthropic 贴身肉搏的当口,这一跤摔得不算轻,但也算给全行业提了个醒:把别人的脸塞进 AI 之前,先问一句”你愿意吗”,大概不该是 optional。

  • 开源AI抢走流量,Anthropic凭什么还在赚大钱

    一个反直觉的现象

    上周,客服 AI 公司 Decagon 的 CEO Jesse Zhang 发了一篇长帖,标题很嚣张:”大家都搞错了企业里的开源 AI”。他抛出一个挺有意思的矛盾:在自己公司里,越来越多成熟的业务正在切到更轻量的模型上,可花在那些最贵的前沿模型上的总预算,几乎一点没少。

    这跟很多人预想的不一样。按常理,开源模型又免费又能自己部署,前沿实验室早该被冲得七零八落。但 Zhang 认为,这俩根本不是对手,更像是同一个生命周期里的两个阶段。

    前沿实验室会继续握着”发现”,开源则会一点点吃掉”生产”。

    用大白话说就是:先用贵的前沿模型把一个新场景跑通、验证它能赚钱,等路子熟了,再挪到便宜的开源模型上规模化。前端探路,后端量产,各司其职。

    数据摆出来,前沿真没吃亏

    Zhang 没给太多数据,但这种数据其实到处都是。Vercel 的 AI 网关看板显示,就这一周,DeepSeek 的 Token 量已经冲到全平台的三分之一以上,背后是 GLM-5.2 的 Z.ai 也挤进了第四。可往下滑到”总支出”那栏,Anthropic 一家还是占了半壁江山。

    OpenRouter 的故事也差不多。DeepSeek V4 Flash 每周处理 5.3 万亿 Token,是最忙的;而最火的前沿模型 Opus 4.8 才 2 万亿出头。但 Opus 每百万 Token 要 1.37 美元,V4 Flash 只要 6 美分,贵了快 23 倍。算下来,花钱的大头还是落在前沿模型手里。连英伟达刚放的 Nemotron 都还没算进去。

    开源AI与前沿模型的双层经济
    开源模型吃下了流量,前沿模型守住了利润(图:TechCrunch)

    为什么开源赢不了”钱”

    一种解释是,AI 能干的活儿增长得太快了,前沿模型只要卡住早期那些最难、最值钱的部署,就能一直待在牌桌上。另一种解释更实在:很多场景就是太难,便宜模型暂时顶不上。

    • 垂直类 AI 公司确实在往轻量模型迁移,这条预言已经应验;
    • 但”GPT 套壳”类创业公司的账,算下来还是稳的;
    • 最关键的是,按 Token 算,前沿厂商死死攥住了最值钱的”溢价 Token”价格。

    作者想起去年九月自己写过的一个比方:基础模型公司到最后可能变成给星巴克供咖啡豆的——自己是商品原料,利润全被应用层赚走。现在看,这个预言只兑现了一半。


    所以眼下的格局,很可能不是”开源干掉前沿”,而是两套玩法长期并存:前沿负责开荒,开源负责量产。只要 AI 能做的事情还在指数级扩张,两边就都还有饭吃。至于这个平衡能维持多久,大概得看下一个更难啃的场景,到底落在谁手里。

  • DeepMind CEO呼吁设立独立机构,监管最强AI模型发布

    给最强模型找个”上市前审查员”

    前沿 AI 模型发布前,到底谁来把关?Google DeepMind 的 CEO 戴密斯·哈萨比斯这周在 X 上发了一篇长文,抛出了一个相当具体的答案:参照金融业的 FINRA,成立一个独立的”标准机构”,在最强模型向公众开放之前,先替大家把危险试出来。

    哈萨比斯把这篇框架文章命名为《前沿 AI 的框架,与一个新时代的开端》。他的设想是:最初,做前沿模型的实验室可以自愿在发布前最多 30 天,把模型交给这个标准机构做评测;一旦评估流程被证明靠谱,就会顺势变成强制要求——通不过,就不能在美国市场部署。发布之后要是冒出严重漏洞,实验室也得配合机构一起修。

    DeepMind CEO Demis Hassabis
    Google DeepMind CEO 戴密斯·哈萨比斯(图源:TechCrunch / Getty Images)

    “这个办法的好处是,它技术上聚焦,同时又不压制创新,还能奖励负责任的做法。”哈萨比斯坦言,机构要跟得上领域加速,并能在风险变大时加码。

    绕开”AI 版 FDA”的政治死结

    AI 监管在美国至今是个烫手山芋。白宫 AI 顾问、a16z 合伙人 Sriram Krishnan 前不久就把话挑明了:行政体系里”不会有 AI 的 FDA”。把标准机构设计成 FINRA 那样的自监管组织,恰恰是给这个分歧留了条缝——它既有约束力,又不用新设一个联邦部门。

    按哈萨比斯的构想,这个机构由开源代表和业内技术专家组成,资金来自各家 AI 实验室,还可以把某些评测外包给专门盯特定风险的 AI 安全小组。测试重点也很明确:网络攻击、生物威胁,以及模型试图绕开自身护栏的”欺骗”行为。

    到底是真能管住最强模型,还是最后变成大玩家们学会”应付考试”的又一处场子,现在谁也给不了答案。但至少在”模型上线前该有人先看一眼”这件事上,行业里最有分量的那几位,似乎正在慢慢达成共识。

  • OpenAI首款硬件曝光:一台会自己动的无屏AI伴侣音箱

    一台会自己动的”伙伴”

    OpenAI 一直说想做硬件,但到底长什么样,外界猜了快一年。这周彭博给出了一版听起来相当具体的描述:它的第一台设备,很可能是一个没有屏幕、能自己在屋里挪动的智能音箱,内部定位是”住在你家里的、像人一样的 AI 伴侣”。

    据彭博周二报道,这台设备还在开发中,被叫做 ChatGPT 的”实体化身”。它不是一个普通智能音箱的换皮——消息人士说它有自己的”性格”,会随着时间主动去了解主人,接入你的邮件之类的数字生活,提供的服务越来越个性化。最出戏的一处描述是,它带有”能自行运动的机械部件”,设计目标就是让人觉得它像个伙伴,而不只是一台机器。

    OpenAI 首款 AI 硬件设备设想图
    OpenAI 首款硬件设备的设想图(图源:TechCrunch / Getty Images)

    一边打官司,一边把第一款硬件的轮廓抛出来,OpenAI 这一步走得相当大胆。它想证明自己不只是个 App,而是要把 ChatGPT 从屏幕里搬进客厅。

    苹果在门外,资本在门外

    负责这件事的团队里,有不少是从苹果出来的老将,当年 iPhone 和 Mac 就是他们参与打造的。这个履历现在反而成了麻烦:就在上周,苹果正式起诉 OpenAI 窃取商业机密,并且放话称目前的指控”只是冰山一角”。OpenAI 当然否认了。彭博援引知情人士称,OpenAI 内部认为新产品和苹果市面上的任何东西差异都很大,”不太可能侵犯苹果的商业机密”。

    OpenAI 不是唯一盯上这块的人。由 Brett Adcock 创立的 AI 实验室 Hark,今年 5 月就拿到了超额认购的 7 亿美元 A 轮,估值冲到 60 亿美元,要做的是所谓”个人智能”——自研模型配定制硬件,做成人和机器之间的”通用接口”。连产品长什么样都还没公布,资本已经先涌进来了。

    说到底,这款会动的音箱能不能成,取决于它能不能真的像一个”伙伴”那样有用,而不是又一个摆在架子上吃灰的电子产品。但 OpenAI 显然已经下定决心,要把赌注押在”无屏、随身、有温度”的硬件路线上。