Sign in

luokai

@luok.ai
2.6K followers 2.5K following 5K posts

For more AI&Tech content, check here www.luok.ai 🍎Apple Die Hard Fan| 苹果骨灰粉 🤖GenAI Observer | GenAI观察者 👨🏻‍🎤Cutting Edge Tech Enthusiast | 科技爱好者

PostsRepliesMedia
luokai @luok.ai · 12/02/2026
jimeng.jianying.com
jimeng.jianying.com
即梦AI - 即刻造梦
即梦AI一站式智能创作平台,即刻造梦。提供AI绘画和AIGC视频创作体验,拥有激发无限创作灵感的社区。让即梦AI开启您的智能创作之旅,探索梦境实现的无限可能!
080
luokai @luok.ai · 12/02/2026
YES
170
luokai @luok.ai · 11/02/2026
But for this project, I generated the clips first and then re-generated them to add lip-sync.
040
luokai @luok.ai · 11/02/2026
This multimodal reference capability is quite rare among current AI video tools. In theory, I could have directly provided the model with edited music or voice clips along with reference images for generation.
140
luokai @luok.ai · 11/02/2026
However, Seedance supports up to 9 images, 3 video clips, and 3 audio clips as reference materials simultaneously for each generated segment.
110
luokai @luok.ai · 11/02/2026
This was a habitual mistake I made while working on this video. Initially, I followed the traditional workflow for video models: first generating reference images, then describing the actions, and so on.
110
luokai @luok.ai · 11/02/2026
After generating the clips, I edited them by adding lip-sync, syncing them with the music, and adjusting the speed of some segments to match the beat.
110
luokai @luok.ai · 11/02/2026
Seedance 2 automatically designs camera angles based on the content, though you can also specify camera movements precisely. In the raw clip below, I didn’t describe camera angles—you can compare it with my final video.
110
luokai @luok.ai · 11/02/2026
1. Overall atmosphere description 2. Key actions 3. Scene description: starting pose, mid-sequence body/hand movements over time, and ending pose 4. Dialogue/lyrics/sound effects at specific timestamps
110
luokai @luok.ai · 11/02/2026
To clarify, I didn’t use any real human dance footage as reference for this video—everything was generated and then edited together. Each segment of my video is based on prompts that generally include the following elements:
190
luokai @luok.ai · 10/02/2026
Every chant, every breath, every siren hit pulses like a declaration of control. It’s not about dancing to the rhythm — it’s about being the rhythm. Minimal. Hypnotic. Absolute. youtu.be/rxWNmzQpW2c
youtu.be
OWN THE BEAT
YouTube video by LUOKAI
030
luokai @luok.ai · 10/02/2026
🔥 When rhythm takes over, power isn’t shown — it’s felt. OWN THE BEAT is raw Brazilian Funk stripped to its essence — no melody, just command.
100
luokai @luok.ai · 10/02/2026
In the past, producing a video like this would have taken me at least a week, and the quality wouldn’t have been nearly as good. Hollywood really needs to start rethinking its approach to content creation.
100
luokai @luok.ai · 10/02/2026
The Seedance 2 model is incredibly powerful, completely overshadowing all other models. This is an original video I created in just one day, though the music was previously made using Suno.
350
luokai @luok.ai · 12/01/2026
This will drive upcoming Apple Intelligence features—including a more personalized Siri—while Apple continues to leverage on-device and Private Cloud Compute to maintain its industry-leading privacy standards.
030
luokai @luok.ai · 12/01/2026
Finally, it’s official: Apple’s next AI leap is… built on Google’s Gemini. 🤯 Apple and Google have signed a multi-year agreement: future Apple Foundation Models will be based on Gemini models and Google Cloud technology.
331
luokai @luok.ai · 10/01/2026
Open-source foundation. Dev-focused sample from Oculus DevTech. Fork it, swap languages, tune models, and build your own MR learning experiences. It’s a baseline to prototype commercial-grade features without starting from zero. Github: github.com/oculus-sampl...
github.com
GitHub - oculus-samples/Unity-SpatialLingo: Spatial Lingo is an open source Unity app for Meta Quest that helps users practice languages through real-world object recognition. Built with Meta SDKs, it...
Spatial Lingo is an open source Unity app for Meta Quest that helps users practice languages through real-world object recognition. Built with Meta SDKs, it’s a template for mixed reality experienc...
030
luokai @luok.ai · 10/01/2026
MR-first UX via Passthrough. You’re learning in your actual environment, not a cartoon room. Roomscale + Hand Tracking + Voice = hands-free practice.
100
luokai @luok.ai · 10/01/2026
It identifies chairs, desks, and more, then overlays nouns/adjectives in your target language. The app listens and judges pronunciation strictly. That’s useful for serious practice, even if it feels tough. Expect real-time feedback and progression into a “final level” with sharper visuals.
100
luokai @luok.ai · 10/01/2026
Built for Meta Quest Passthrough, it detects objects around you, overlays translated words, and listens as you speak. A playful 3D guide gives real-time pronunciation feedback, turning your room into a dynamic classroom. It’s positioned as an open-source challenger to commercial MR language apps.
meta.com
Spatial Lingo: Language Practice on Meta Quest
Spatial Lingo is an open source showcase app for Meta Quest that transforms your space into an interactive language practicing playground. Instantly identify and translate real-world objects, practice...
100
luokai @luok.ai · 10/01/2026
A Meta Quest open-source MR app turns your room into a language lab. Spatial Lingo shows how mixed reality + AI can teach vocab by labeling your real world—now open-source.
240
luokai @luok.ai · 10/01/2026
sref: style reference control. Use sref to steer aesthetic toward a target look while keeping your prompt. Handy for series consistency, brand vibes, or matching a particular artist’s feel.
010
luokai @luok.ai · 10/01/2026
Prompt following for specifics. Niji 7 improves on complex, multi‑clause requests. It’s more literal with ordering and constraints, so you can stack attributes without losing key elements.
100
luokai @luok.ai · 10/01/2026
Coherency: “what you ask is what you get.” Better compliance with spatial cues (left/right), colors, counts. E.g., “red cube left, blue cube right” renders correctly more often, cutting prompt wrangling.
100
luokai @luok.ai · 10/01/2026
Core: “Crystal Clarity.” Sharper reflections and eye details reduce muddiness in faces and highlights. Expect fewer artifacts in glossy surfaces and more readable micro‑features—think eyelashes, irises, jewelry.
100
luokai @luok.ai · 10/01/2026
Key stats: Coherency: major improvement vs prior Niji Prompt following: stricter left/right, color, object placement Compatibility: backwards support incl. –sv 4; use –niji 7 in Discord or “Version: Niji 7” on web
100
luokai @luok.ai · 10/01/2026
Niji 7 just landed. The latest Niji focuses on sharper eyes, tighter coherency, and better prompt adherence. It keeps legacy flags and adds sref tweaks for style control. After 18 months of training, this release targets fewer misses and more faithful outputs for anime creators.
130
luokai @luok.ai · 09/01/2026
I’ve connected with LuxReal and got three redeem codes for you to try more. Share your test results and videos in the comments—first three get the codes via DM. Try it now: www.luxreal.ai
luxreal.ai
LuxReal | Al-Powered Creator for Product Videos
LuxReal creates high-quality product videos across beauty, electronics, PMCG, toys, food & beverage, and more. With just a single image, LuxReal generates cinematic, consistent product ads—delivering ...
000
luokai @luok.ai · 09/01/2026
The next step for AI video isn’t about being more “flashy,” but more “stable.” LuxReal’s approach is still in its early stages, but the direction is right. Below is the link—feel free to join the beta test. Share the product ads you create with LuxReal and let me know about your experience.
110
luokai @luok.ai · 09/01/2026
If AI video is to truly become a “tool,” I lean toward this path: first, ensure the video makes sense in a 3D world, then focus on style and flair. Controllability, reusability, and credibility—these all stem from spatiotemporal consistency.
110
luokai @luok.ai · 09/01/2026
- Multi-agent collaboration: Instead of a closed approach, it integrates industry-leading algorithms, emphasizing “multiple capabilities working together to produce a stable video.”
100
luokai @luok.ai · 09/01/2026
Two aspects of their underlying tech resonate with me: - 3D generative model (Lux3D): Converts text/images into 3D representations, making materials and volumes more realistic.
100
luokai @luok.ai · 09/01/2026
- Choose the subject: A person or product—just upload an image. - Set the stage: Environment, actions, lighting. - Define the logic: Camera movement and subject motion. The steps aren’t complicated, but the core lies in how “stability” is achieved at the foundational level.
100
luokai @luok.ai · 09/01/2026
LuxReal’s approach starts with 3D and integrates multiple capabilities to solve the “stability” problem. Their workflow is straightforward (which is why I think it’s practical):
100
luokai @luok.ai · 09/01/2026
Stable, 3D-aware “spatiotemporal consistency” so objects, people, and scenes remain coherent across frames. LuxReal is laying the first brick on this path with a 3D-first, multi‑agent approach. Try it now: www.luxreal.ai At the end of the post, I have some free redeem codes for you to try.
100
luokai @luok.ai · 09/01/2026
AI video is stuck in “fun”—the next leap is making it truly useful. Over two years since Sora’s splashy debut, generative video is dazzling but still struggles to land in everyday workflows. The missing piece?
130
luokai @luok.ai · 09/01/2026
Core idea: two‑stage monocular data. Static, texture‑rich images supervise Gaussian attributes; dynamic motion sequences optimize pose‑dependent deformation and lighting changes—bridging single‑view limitations with tailored signals. Project: acennr-engine.github.io/HRM2Avatar/
acennr-engine.github.io
HRM^2Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone Scans
We introduce HRM^2Avatar, a framework that produces photorealistic, personalized avatars from phone scans, supporting real-time rendering and animation on mobile devices. Combining dual-phase capture,...
000
luokai @luok.ai · 09/01/2026
A fully GPU pipeline pushes sub‑frame responsiveness for AR/VR and social apps, making studio‑grade detail accessible. Key stats iPhone 15 Pro Max: 2048×945 120FPS Apple Vision Pro: 1920×1824×2 90FPS 533,695 splats; 5‑minute single‑phone capture
100
luokai @luok.ai · 09/01/2026
HRM^2Avatar turns a single iPhone scan into a high‑fidelity, real‑time digital human—on mobile. ⚡ It combines mesh‑driven clothing deformation with illumination‑aware Gaussians, learning textures from static shots and dynamics from motion clips.
130
luokai @luok.ai · 09/01/2026
— Hard Surface Boost: Crisper edges and tighter structure for mechanical + geometric models. — Low Poly Mode: Purpose-built for game devs; efficient wireframes optimized for real-time.
030
luokai @luok.ai · 09/01/2026
Meshy 6 just raised the ceiling for AI 3D characters—cleaner topology, sharper forms, truer anatomy. Upgrades that you’ll feel in production: — Refined Geometry: Cleaner, anatomically correct meshes that deform better when rigged.
150
luokai @luok.ai · 15/12/2025
With VIDEO O1, you can define Start & End Frames between 3–10s for smoother transitions and precise timing. From punchy, high-impact beats to immersive cinematic moves, your story flows exactly as intended. Plus, there’s a 720p mode that keeps all the 1080p features for faster renders.
010
luokai @luok.ai · 15/12/2025
Kling just gave video creators granular control over pacing—start/end frames with selectable durations. Key stats Start/End Frame duration: 3–10s Resolution modes: 1080p and 720p (same features)
150
luokai @luok.ai · 05/12/2025
It’s currently gated to AI Ultra subscribers in the Gemini app, with usage under the “Thinking” model + prompt bar toggle. Key stats Access: AI Ultra only Modality: iterative, multi-hypothesis Rollout: app-first
030
luokai @luok.ai · 05/12/2025
Gemini 3’s new Deep Think mode tackles problems by exploring multiple hypotheses at once ⚡ Positioned as Google’s most advanced reasoning tier, Deep Think runs iterative rounds to refine outputs—especially for complex tasks like code visualization, prototyping, and nuanced analysis.
181
luokai @luok.ai · 05/12/2025
Core idea: Max expressions = believable characters Richer facial cues and delivery land humor, urgency, and sincerity. This moves avatars from “demo” to “cast member” for social and ads. Example: a spokesperson that can smile, pause, and emphasize like a human.
000
luokai @luok.ai · 05/12/2025
Core idea: Full‑length avatar performances 5 minutes means complete narratives—intro, arc, CTA—without stitching multiple clips. Think product explainers or mini‑ads in one take, with consistent character and emotion. 🎬
100
luokai @luok.ai · 05/12/2025
For 12 hours only, they’re rewarding engagement with credits and picking 200 winners for a 1‑month Standard Plan, delivered via DM. Techy, playful, and very creator‑friendly.
100
luokai @luok.ai · 05/12/2025
Kling AI just leveled up avatars to full 5‑minute performances. Wild. Avatar 2.0 is upgraded and expressive—built to handle explainers, ads, songs, and stories like real characters. Key stats 5‑minute avatar acts - Max expressions
170
luokai @luok.ai · 05/12/2025
The T800 packs 29 DOF, millisecond‑level sensor fusion, and active leg cooling to sustain high‑intensity operation across logistics, services, and human‑robot collaboration. Key stats 173 cm | 75 kg | 29 DOF + 7 DOF/hand | 3 m/s | up to 4 hrs | 450 N·m torque | AGX Orin 64G, 275 TOPS | 360° LiDAR
youtube.com
EngineAI T800 BTS Footage: Setting the Record Straight on CGI Rumors
YouTube video by Engineai Robot
000