meta-llama/Llama-3.1-8B-Instruct
New text-generation model. Tags: llama, text-generation, facebook, meta, pytorch, llama-3, conversational, en
Buscar
868
New text-generation model. Tags: llama, text-generation, facebook, meta, pytorch, llama-3, conversational, en
New image-text-to-text model. Tags: gguf, unsloth, fine tune, heretic, uncensored, abliterated, ara, MTP GGUF Quants
Assine o Pro, sem anúncios
Faça upgrade para uma leitura sem interrupções e acesso prioritário a novas fontes.
Ver preçosNew text-generation model. Tags: mlx, qwen3_5_moe, moe, edge-inference, prerouter, lora, ssd-offload, text-generation
New text-to-speech model. Tags: audio, speech, text-to-speech, zero-shot-tts, voice-cloning, speech-generation, speech-editing, speech-enhancement
TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model in Amazon Bedrock Knowledge Bases, bringing fully managed natural language search to video, image, and audio content. This walkthrough shows how to build a knowledge base powered by Marengo 3.0 and run semantic queries against your media.
Speech large language models (SpeechLLMs) offer reduced latency and retain paralinguistic nuances that are typically lost in cascaded automatic speech recognition (ASR) and text-based LM architectures. However, they continue to lag behind text-only LLMs on complex reasoning tasks, while real-time spoken interaction imposes strict latency constraints. Although prior works employ Chain-of-Thought (CoT) and concurrent reasoning to enhance reasoning capabilities without inducing prohibitive delays, an inherent accuracy-latency trade-off persists. In this paper, we investigate whether a streaming SpeechLLM can dynamically revise its reasoning traces on the fly. We introduce RetroThinker, a multi-stage post-training framework that equips the Moshi model to self-verify and forward-correct CoT steps during inference. RetroThinker combines supervised fine-tuning (SFT) on curated retrospective thinking data with length-based direct preference optimization (DPO) to optimize retrospective during early reasoning (i.e., reasoning concurrently while the user speaks). Evaluated on the GSM8K benchmark, RetroThinker significantly improves the accuracy-latency trade-off over non-retrospective baselines, achieving an 11% absolute accuracy gain at a comparable latency.
New text-to-audio model. Tags: yue2, music-generation, symbolic-planning, agentic-editing, custom_code, text-to-audio, zh, en
New text-to-image model. Tags: diffusers, image-generation, image-editing, image-to-image, text-to-image, en, zh, arxiv:2609.03796
Commercial text-to-image systems silently revise user prompts before generating images, a step users typically cannot disable or even see. Yet, existing audits of cultural bias examine only the final images and treat generation as a single pipeline, so they cannot tell where the bias originates. We introduce WORLDVIEW, a multilingual benchmark of 8,960 prompts across 15 languages and 31 language-context pairings. Using it, we audit the revision layer in three systems (DALL-E-3, Imagen-4, GPT-Image-1.5) through a three-step analysis of how heavily it marks each cultural context, whether it flattens that context into a narrow vocabulary, and whether that vocabulary is stereotypical. Relative to a no-context English baseline, the US is the least-marked context, while non-Western and non-Anglophone contexts are marked far more heavily, flattened into narrow vocabularies applied across topically diverse prompts, and reduced to recognizable cultural stereotypes. Comparing images from original versus revised prompts on models without a revision layer, we identify the layer itself as a previously undocumented, causal source of this stereotyping. To locate cultural bias, and fix it, we must audit the system as deployed, not the model alone.
New text-to-image model. Tags: diffusers, text-to-image, image-generation, flux, en, license:other, diffusers:FluxPipeline
New text-generation model. Tags: k2_horizon, text-generation, k2-horizon, 375b, moe, open-weights, ifm, conversational
New text-generation model. Tags: gguf, minicpm, minicpm5, llama, text-generation, long-context, tool-calling, on-device
New text-generation model. Tags: bailing_hybrid, finance, financial-research, agents, tool-use, long-context, mixture-of-experts, text-generation
New text-generation model. Tags: gguf, llama.cpp, ollama, lm-studio, sparkx2_5, text-generation, en, zh
New text-generation model. Tags: llama, text-generation, minicpm, minicpm5, long-context, tool-calling, on-device, edge-ai
New text-to-speech model. Tags: text-to-speech, base_model:Aratako/Irodori-TTS-v4.1-Small, base_model:finetune:Aratako/Irodori-TTS-v4.1-Small, license:mit
New text-to-image model. Tags: diffusers, text-to-image, en, base_model:krea/Krea-2-Raw, base_model:finetune:krea/Krea-2-Raw, license:other, diffusers:Krea2Pipeline